PyTorch Multi-Process Inference Weight Sharing Via Inter-Process Communication 07-31-2026 09-02-2026 blog 17 minutes read (About 2564 words)Avoiding Weight Duplication In PyTorch Multi-Process Inference PyTorch, CUDA, AOTInductor, Deep Learning Inference Read More
Synchronizations With TorchRec KeyedJaggedTensor 06-05-2026 06-05-2026 blog 8 minutes read (About 1188 words)Efficiently Using TorchRec KeyedJaggedTensor In GPU Systems PyTorch, GPU, Deep Learning Inference, TorchRec Read More
PyTorch AOTInductor Hybrid Lowering 05-28-2026 05-28-2026 blog 8 minutes read (About 1224 words)A Hybrid DeviceExecution Inference Engine from PyTorch PyTorch, Deep Learning Inference Read More
PyTorch Triton Kernel Transparent Tracing and Compilation 05-22-2026 05-22-2026 blog 27 minutes read (About 4044 words)Exporting and Compiling Triton Kernels for AOTInductor PyTorch, AOTInductor, Deep Learning Inference, Triton Read More
PyTorch Fake Export 05-17-2026 05-17-2026 blog 11 minutes read (About 1634 words)Eliminating PyTorch Model Data Memory Allocation for PyTorch Export Verification PyTorch, Deep Learning Inference Read More
How To Debug Deep Learning Inference Applications 01-01-2024 01-01-2024 article 23 minutes read (About 3511 words)First Principles of Evaluating Deep Learning Inference Deep Learning, Software Engineering, Numerical Errors, Deep Learning Inference Read More