Lei Mao's Log Book
Lei Mao's Log BookCurriculumBlogArticlesProjectsPublicationsReadingsLifeEssayPhotographyArchivesCategoriesTagsFAQs
  • Tags
  • CUDA

CUDA Graph In The Context of Multi-Stream Execution

 08-20-2026 08-20-2026 blog 4 minutes read (About 556 words)
When CUDA Graph Does Not Quite Work As Expected

 
CUDA, 
CUDA Graph  
  Read More

PyTorch Asynchronous Assert

 08-14-2026 08-14-2026 blog 11 minutes read (About 1587 words)
Using torch._assert_async for PyTorch CUDA Model Development

 
CUDA, 
PyTorch, 
Perfetto, 
Triton  
  Read More

Predicated Execution VS Conditional Execution

 07-01-2026 07-02-2026 blog 17 minutes read (About 2609 words)
where VS cond

 
Accelerated Computing, 
CUDA, 
TensorRT, 
PyTorch, 
GPU, 
AOTInductor, 
TorchInductor, 
JAX, 
XLA, 
TPU  
  Read More

PyTorch Custom Operation

 05-10-2026 05-10-2026 blog 23 minutes read (About 3501 words)
Implementing PyTorch Custom Operations In C++ and CUDA Using torch.library

 
CPP, 
Python, 
CUDA, 
PyTorch, 
AOTInductor  
  Read More

How Is FARS, The Fully Automated Research System?

 04-22-2026 04-22-2026 blog 5 minutes read (About 740 words)
The AI Just Tried To Fool People

 
Artificial Intelligence, 
CUDA, 
Research, 
CUDA Graphs, 
Mixture-of-Experts  
  Read More

Page Table for Page-Locked Host Memory

 04-12-2026 04-12-2026 blog 17 minutes read (About 2541 words)
Page Table GPU Memory Overhead and Sharing Page-Locked Host Memory Across Processes

 
CUDA, 
NVIDIA, 
Computer Architecture, 
GPU, 
Memory Management  
  Read More

CUDA_LAUNCH_BLOCKING=1

 03-20-2026 03-20-2026 blog 4 minutes read (About 599 words)
Debugging CUDA Applications

 
CUDA, 
Debug  
  Read More

CUDA Shared Memory Bank Conflict-Free Vectorized Access

 02-13-2026 02-13-2026 blog 14 minutes read (About 2060 words)
Instruction-Level Phase Based Bank Conflict-Free Execution

 
CUDA, 
NVIDIA, 
Parallel Computing, 
GPU  
  Read More

CUDA Rendezvous Stream

 01-26-2026 01-26-2026 blog 11 minutes read (About 1690 words)
Simplifying Synchronization Complexities Using CUDA Rendezvous Streams

 
CUDA, 
NVIDIA, 
Parallel Computing, 
GPU  
  Read More

PyTorch CUDA Graph Capture

 01-12-2026 08-07-2026 blog 23 minutes read (About 3479 words)
Using PyTorch CUDA Graph APIs

 
CUDA, 
PyTorch, 
CUDA Graph, 
Perfetto  
  Read More

NVIDIA NVML GPU Statistics

 12-25-2025 12-25-2025 blog 15 minutes read (About 2214 words)
Mimicking nvidia-smi dmon Using NVIDIA NVML

 
CPP, 
CUDA, 
NVIDIA, 
GPU, 
NVML  
  Read More

NVIDIA Tensor Core TN Layout MMA Instruction

 12-06-2025 12-06-2025 blog 16 minutes read (About 2389 words)
GEMM Layout, History, Performance, and Implementation

 
CPP, 
CUDA, 
NVIDIA, 
CUTLASS, 
CuTe, 
MMA, 
GEMM, 
Tensor Core  
  Read More
Previous
Next
  • 1
  • 2
  • …
  • 7
Lei Mao

Lei Mao

Artificial Intelligence Machine Learning Computer Science

Menlo Park, California

Posts

1433

Categories

8

Tags

834

  Follow   Sponsor

Advertisement


Categories

  • article21
  • blog585
  • essay362
  • life342
  • miscellaneous2
  • photography93
  • project20
  • reading8

follow.it

Recents

08-27-2026

AOTInductor External Weight Storage and Weight Streaming Update

blog

08-26-2026

宁波特产伴手礼

essay

08-23-2026

Edgewood Park Natural Preserve 徒步

life

08-23-2026

Edgewood Park Natural Preserve

photography

08-22-2026

2026 Race2Unravel 10K 竞赛

life

Archives

  • August 202620
  • July 202623
  • June 202621
  • May 202624
  • April 202618
  • See All >>

Tags

Outdoors347
California279
Hiking259
CPP122
Photography109
Mathematics103
Deep Learning87
Wildlife83
Running80
CUDA78
Bird77
Racing56
Movie45
Python38
Software Engineering36
Machine Learning35
China34
Linux33
NVIDIA32
Statistics32
See All >>
Lei Mao's Log Book

© 2017-2026 Lei Mao  Powered by Hexo & Icarus
Site UV:  Site PV:

×