A critical assessment of the ~$5M French seed promising “30x faster” LLM inference on standard GPUs — a headline number Kog’s own website quietly walks down to 3.5x, backed by zero independent benchmarks, and structurally exposed to the free first-party stacks of the very chipmakers it depends on.
The 30x derives from a batch-size-1, 2B custom coding model, measured against a theoretical upper bound Kog itself calls “not guaranteed achievable.” The only real-model figure Kog publishes — 1,368 tok/s on Llama-3 8B — and its own site’s “up to 3.5x” imply a far smaller multiplier. No independent or head-to-head benchmark exists.
A single-kernel, hardware-specific inference engine is precisely the latency margin Nvidia (TensorRT-LLM) and AMD (ROCm/vLLM) are motivated to absorb into their free, bundled stacks. Kog sits between those free first-party tools and better-funded neutral rivals (ZML $20M, Infinity $15M) — on ~$5M.
The August 2026 press appears to be the media surfacing of the same seed already tied to Kog’s October 2025 French Tech 2030 label — not a fresh raise. The CEO says he still expects to raise a Series A only after a September milestone, meaning as of the announcement Kog is still on its original seed capital.
Key Finding: Kog is a credibly-staffed, well-timed, thinly-capitalized deep-tech bet whose signature “30x” is an unverified batch-1 theoretical ceiling that its own website walks down to “3.5x.” The engineering looks real — open-sourced demo model, unusually honest blog caveats — but every performance number traces back to Kog itself, there is no named paying customer, and the frontier-scale demo that gates the Series A has not shipped.
The headline number is engineered from the single most flattering configuration possible. Follow the qualifiers and the multiplier shrinks at every step.
The marketing headline on kog.ai. No model, batch size, or baseline attached to the number itself.
Measured at single-request decoding — the config that most flatters a memory-streaming engine and least reflects production economics.
The demo ran on Laneformer 2B, Kog’s own purpose-built coding model (~50% HumanEval) — not a standard frontier model.
3,000 tok/s on 8×MI300X — which Kog’s own blog labels an “upper bound, not guaranteed achievable,” at just 36% memory-bandwidth utilization.
The number Kog’s own site cites for realistic AMD Instinct workloads. Its only real-model figure: 1,368 tok/s on Llama-3 8B.
No independent lab, no head-to-head vs. vLLM / TensorRT-LLM / SGLang, no named customer. Every number is self-reported.
To its credit, Kog’s engineering candor is real. The company open-sourced its demo model and its blog explicitly lists what was excluded from the benchmark — no quantization, no speculative decoding, no pruning, no KV-cache compression — and calls the figures upper bounds. The problem is not dishonesty in the fine print; it is a headline number chosen to be ~10x larger than the fine print the company itself publishes.
Kog’s entire value proposition is defensible only as long as Nvidia (TensorRT-LLM), AMD (ROCm/vLLM), and open-source SGLang leave a single-kernel latency margin on the table. A hardware-specific inference engine is exactly the kind of optimization the chipmakers are motivated to absorb into their free, bundled software — margin that evaporates into the vendors’ stacks the moment they prioritize it. Kog has never publicly explained how an 11-person team sustains a moat against the very silicon vendors it depends on.
Kog is wedged between free first-party stacks below it and far-better-capitalized neutral rivals beside it — on roughly ~$5M of seed capital.
The closest analog — also French, also chip-agnostic (Nvidia/AMD/TPU/Metal/Intel), bypassing CUDA. Backed by LeCun, Hykes, and Hugging Face; named by TechCrunch as Kog’s direct rival, with 4x the capital.
“Software layer that makes any AI chip inference-ready,” auto-generating low-level kernels — the same wedge as Kog, with OpenAI/Anthropic-researcher backing and a stronger network.
The real threat: free, vendor-backed, improving every release cycle. The single-request latency Kog charges for is theirs to reclaim — and they are motivated to do so.
The inference-cloud incumbents. They validate that inference is the 2026 capital magnet — and set the bar for how much a serious inference company is expected to raise.
The “build new silicon” camp Kog defines itself against. Purpose-built inference hardware attacking the same latency problem from the metal up.
Barely enough to reach the September Series-A-gating demo, with no margin for error — a severe capital asymmetry against every player above.
Market timing is the strongest part of the thesis. “More performance from existing GPUs” is squarely in-thesis for the 2026 inference boom. But timing cuts both ways: the same boom has funded neutral rivals 3–4x deeper, and the incumbents whose silicon Kog optimizes are the ones best positioned to close the gap for free.
Seven structural concerns, each stated as a falsifiable claim.
The 30x rests on a batch-1, 2B custom-model, theoretical-ceiling comparison; the site itself shows “up to 3.5x.” Falsifiable: an independent Llama-3 8B run vs. optimized vLLM/TensorRT-LLM.
A single-kernel, GPU-specific engine is exactly what Nvidia and AMD can absorb for free. Falsifiable: watch whether their next release ships comparable single-request latency.
All numbers self-reported; “200 leads” and unnamed “design partners” are not revenue. Falsifiable: a named, referenceable customer running a frontier model in production.
The entire Series A thesis hinges on a September “10x on a major model” demo that does not yet exist. Falsifiable: does it land, and at what memory-bandwidth utilization?
Real inference cost is throughput-at-concurrency; Kog publishes no cost-per-token or high-batch numbers. Falsifiable: concurrent-throughput and $/1M-token disclosures.
No ROCm/CUDA version, firmware, interconnect, or Scaleway-specific requirements published for the 8×GPU results. Falsifiable: reproducibility on a non-Scaleway MI300X node.
11 people, with capital and label tied to French/EU deep-tech policy; the thesis partly rests on compute-sovereignty politics rather than open-market performance wins. Falsifiable: commercial traction outside French/EU-subsidized channels.
Kog is a real team, well-timed, taking the hardest way in. The engineering is credible and the candor in its technical blog is unusual. But the signature “30x” is an unverified batch-1 theoretical ceiling that Kog’s own website walks down to “3.5x,” the ~$5M seed is secondary-sourced and appears to be the same round as the October 2025 French Tech 2030 label — not a fresh raise — and the moat question is the one Kog has never answered: it sells the exact latency margin that Nvidia’s and AMD’s free first-party stacks are motivated to reclaim. The entire Series A thesis rests on a September milestone that has not shipped.
Based entirely on publicly available information, including the TechCrunch report of August 14, 2026. Every performance figure in this report is self-reported by Kog unless otherwise noted; the ~$5M round size is confirmed only via secondary aggregation, not a primary filing.