Skip to main content
Back to News
news/AI Infrastructure

French Startup Kog Squeezes More AI Inference Out of GPUs

French startup Kog claims software optimization can unlock far faster LLM inference on standard GPUs like AMD MI300X and Nvidia H200, without new hardware.

Stefan Trbojevic

Stefan Trbojevic

15 August 20262 min read
LinkedIn
French startup Kog squeezing more AI inference out of datacenter GPUs

The takeaway

Software-level optimization, not new silicon, may be the fastest route to cheaper, lower-latency AI inference.

Why it matters for builders

For builders running agentic workloads, Kog's bet signals that GPU inference is far from optimized. Software-level acceleration — monokernels, delayed tensor parallelism, and continuous weight streaming — could unlock 10x-plus speedups on hardware teams already own, cutting token costs and latency for real-time agents without a hardware upgrade.

French Startup Kog Squeezes More AI Inference Out of GPUs

The race for faster AI inference has a new contrarian entrant. Markets rewarded Cerebras and its purpose-built chips with a warm IPO debut in May, but French startup Kog is betting that conventional GPUs still have far more performance locked inside them than anyone has managed to extract.

Software over silicon

Kog hit the front page of Hacker News in May with a tech preview aimed at proving that "extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own" — namely the AMD MI300X and Nvidia H200. Its Kog Inference Engine (KIE) hit roughly 3,000 output tokens per second on a single request, though the demo ran on a purpose-built 2-billion-parameter model, Laneformer 2B, which the startup has since open-sourced.

Software-level GPU inference optimization unlocks faster decoding on existing datacenter GPUs

That demo translated into real demand. "We had 200 tangible business leads," CEO Gaël Delalleau told TechCrunch. The first target use case is software engineering, where veteran Claude Code users "sometimes have to wait hours to get results."

The 30x promise

Kog's headline ambition is "30x faster LLM inference," but the gap between a 2B-parameter model and frontier-scale LLMs is substantial. The company has refocused entirely on accelerating larger models, since prospective customers aren't willing to fine-tune small ones.

Delalleau rejects the pessimism around GPUs for decoding. "GPUs have a bright future," he said, pointing to ever-increasing memory bandwidth that "only begs to be unlocked." The tradeoff is a labor-intensive approach: Kog's 11-person team spends "several weeks or even months" reverse-engineering each new GPU, a mindset Delalleau traces to his background in solid-state physics and DEFCON-level offensive security.

Why it matters

Kog isn't alone. Fellow French startup ZML released hardware-agnostic software that bypasses Nvidia's CUDA, while Stanford's Hazy Research lab pushes in a similar direction. Kog's deeper, GPU-specific focus is what sets it apart, and it arrives as inference speed becomes the industry's new battleground.

Backed by Bpifrance, French Tech 2030, and Scaleway, Kog is also riding European AI-sovereignty tailwinds. The company plans to demonstrate its first major model running at 10x speed by September — a milestone that would unlock its Series A. For builders watching inference costs climb, that's a signal worth tracking.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

15 August 2026

Updated

15 August 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.