Cohere Megakernel
library Your tags
Your notes
A serving engine that executes an entire decode forward pass of North Mini Code inside one persistent CUDA megakernel on H100 (sm_90a, batch up to 8), reaching 62% of speed-of-light at batch 1 against vLLM's 39%: 1.58× vLLM at batch 1 (292 tokens/s) and 1.25 to 1.41× end to end against vLLM v0.24 out to a 256K context, behind an OpenAI-compatible API with continuous batching, paged KV, and prefix caching. By Xiaochun Tong, Conway Zhu, and Donglu Wang of Cohere's Foundations team; Apache 2.0.
Library
License Apache-2.0