Meta AI Updates: October 8, 2026
1. Meta Rewrites TorchRec’s Table Batched Embedding Kernels in Triton and Beats the CUDA Versions
Meta. A PyTorch blog post explains how Meta reimplemented the Table Batched Embedding (TBE) forward and backward kernels in FBTriton, its Triton fork, replacing the legacy CUDA Jinja templates. TBE handles embedding lookups for recommendation models sharded across thousands of GPUs. Across 307 shard configurations on GB200 (exact row-wise Adagrad, FP16 weights), the Triton forward pass has a median speedup of 1.28x. In the backward pass, raising the cutoff for switching to cooperative kernels from a run length of 32 to 256 makes some shapes 4.3x faster, with Nsight showing 3,948 GB/s of memory throughput against 678 GB/s for CUDA. Blackwell-specific features (cluster launch control, device-scope fences, TMA bulk reductions) are optional flags on a single kernel body that also runs on Hopper and AMD. Triton still loses on 11 shards whose runs are all shorter than 4. The code lives in TorchRec under torchrec/distributed/triton_tbe. Source