Lei Mao's Log Book
Lei Mao's Log BookCurriculumBlogArticlesProjectsPublicationsReadingsLifeEssayPhotographyArchivesCategoriesTagsFAQs
  • Tags
  • CUDA

AOTInductor Triton SASS Inspection

 10-02-2026 10-02-2026 blog 9 minutes read (About 1352 words)
Triton Kernel Behavior Verification In AOTInductor

 
Accelerated Computing, 
CUDA, 
PyTorch, 
AOTInductor, 
Triton, 
JIT  
  Read More

CUDA Device Max Connections

 09-26-2026 09-26-2026 blog 10 minutes read (About 1514 words)
Unlocking GPU Concurrency

 
CPP, 
Accelerated Computing, 
CUDA, 
NVIDIA, 
AMD  
  Read More

CUDA Thread Block Swizzle

 09-20-2026 09-20-2026 blog 36 minutes read (About 5405 words)
Influencing L2 Cache Efficiency Using CUDA Thread Block Swizzle

 
Accelerated Computing, 
CUDA, 
Triton  
  Read More

CUDA Multi-Process Service

 09-08-2026 09-08-2026 blog 8 minutes read (About 1199 words)
Executing CUDA Kernels Concurrently From Multiple Processes

 
Accelerated Computing, 
CUDA, 
PyTorch, 
MPS  
  Read More

CUDA Graph In The Context of Multi-Stream Execution

 08-20-2026 08-20-2026 blog 4 minutes read (About 556 words)
When CUDA Graph Does Not Quite Work As Expected

 
CUDA, 
CUDA Graph  
  Read More

PyTorch Asynchronous Assert

 08-14-2026 08-14-2026 blog 11 minutes read (About 1587 words)
Using torch._assert_async for PyTorch CUDA Model Development

 
CUDA, 
PyTorch, 
Perfetto, 
Triton  
  Read More

PyTorch Multi-Process Inference Weight Sharing Via Inter-Process Communication

 07-31-2026 09-02-2026 blog 17 minutes read (About 2564 words)
Avoiding Weight Duplication In PyTorch Multi-Process Inference

 
Deep Learning Inference, 
CUDA, 
PyTorch, 
AOTInductor  
  Read More

Predicated Execution VS Conditional Execution

 07-01-2026 07-02-2026 blog 17 minutes read (About 2609 words)
where VS cond

 
Accelerated Computing, 
CUDA, 
TensorRT, 
PyTorch, 
GPU, 
AOTInductor, 
TorchInductor, 
JAX, 
XLA, 
TPU  
  Read More

PyTorch Custom Operation

 05-10-2026 05-10-2026 blog 23 minutes read (About 3501 words)
Implementing PyTorch Custom Operations In C++ and CUDA Using torch.library

 
CPP, 
Python, 
CUDA, 
PyTorch, 
AOTInductor  
  Read More

How Is FARS, The Fully Automated Research System?

 04-22-2026 04-22-2026 blog 5 minutes read (About 740 words)
The AI Just Tried To Fool People

 
Artificial Intelligence, 
CUDA, 
Research, 
CUDA Graphs, 
Mixture-of-Experts  
  Read More

Page Table for Page-Locked Host Memory

 04-12-2026 04-12-2026 blog 17 minutes read (About 2541 words)
Page Table GPU Memory Overhead and Sharing Page-Locked Host Memory Across Processes

 
CUDA, 
NVIDIA, 
Computer Architecture, 
GPU, 
Memory Management  
  Read More

CUDA_LAUNCH_BLOCKING=1

 03-20-2026 03-20-2026 blog 4 minutes read (About 599 words)
Debugging CUDA Applications

 
CUDA, 
Debug  
  Read More
Previous
Next
  • 1
  • 2
  • …
  • 7
Lei Mao

Lei Mao

Artificial Intelligence Machine Learning Computer Science

Menlo Park, California

Posts

1468

Categories

8

Tags

845

  Follow   Sponsor

Advertisement


Categories

  • article21
  • blog591
  • essay370
  • life355
  • miscellaneous2
  • photography101
  • project20
  • reading8

follow.it

Recents

10-03-2026

East Bay Regional Park Trails Challenge 2026

essay

10-03-2026

2026 Pace For Peace 10K 竞赛

life

10-03-2026

Sycamore Grove Park 徒步

life

10-03-2026

Sycamore Grove Park

photography

10-02-2026

AOTInductor Triton SASS Inspection

blog

Archives

  • October 20265
  • September 202624
  • August 202625
  • July 202623
  • June 202621
  • See All >>

Tags

Outdoors361
California294
Hiking269
CPP123
Photography117
Mathematics103
Wildlife91
Deep Learning87
Bird85
Running84
CUDA83
Racing60
Movie48
Python38
Software Engineering36
China35
Machine Learning35
PyTorch34
Accelerated Computing33
Linux33
See All >>
Lei Mao's Log Book

© 2017-2026 Lei Mao  Powered by Hexo & Icarus
Site UV:  Site PV:

×