Kog: A $5M Bet That Software Can Out-Run the Chipmakers

A critical assessment of the ~$5M French seed promising “30x faster” LLM inference on standard GPUs — a headline number Kog’s own website quietly walks down to 3.5x, backed by zero independent benchmarks, and structurally exposed to the free first-party stacks of the very chipmakers it depends on.

ProofStory Research August 14, 2026

~$5M Seed Co-Led by Varsity VC — Surfaced August 14, 2026

Paris-based Kog builds the Kog Inference Engine (KIE), a low-level software layer pitched as a drop-in vLLM replacement that rewrites LLM decoding as a memory-streaming problem on standard datacenter GPUs (AMD MI300X, Nvidia H200). Backers: Varsity VC (co-lead), Scaleway, Bpifrance Deep Tech, French Tech 2030.

~$5M
Seed (secondary-sourced)
30x
Claimed max speedup
3.5x
Speedup on its own site
11
Person team (5 PhDs)

Three Core Questions

01

“Is the 30x Actually Real?”

The 30x derives from a batch-size-1, 2B custom coding model, measured against a theoretical upper bound Kog itself calls “not guaranteed achievable.” The only real-model figure Kog publishes — 1,368 tok/s on Llama-3 8B — and its own site’s “up to 3.5x” imply a far smaller multiplier. No independent or head-to-head benchmark exists.

02

“Where Is the Moat?”

A single-kernel, hardware-specific inference engine is precisely the latency margin Nvidia (TensorRT-LLM) and AMD (ROCm/vLLM) are motivated to absorb into their free, bundled stacks. Kog sits between those free first-party tools and better-funded neutral rivals (ZML $20M, Infinity $15M) — on ~$5M.

03

“Is This Even a New Round?”

The August 2026 press appears to be the media surfacing of the same seed already tied to Kog’s October 2025 French Tech 2030 label — not a fresh raise. The CEO says he still expects to raise a Series A only after a September milestone, meaning as of the announcement Kog is still on its original seed capital.

Key Finding: Kog is a credibly-staffed, well-timed, thinly-capitalized deep-tech bet whose signature “30x” is an unverified batch-1 theoretical ceiling that its own website walks down to “3.5x.” The engineering looks real — open-sourced demo model, unusually honest blog caveats — but every performance number traces back to Kog itself, there is no named paying customer, and the frontier-scale demo that gates the Series A has not shipped.

The Numbers

Founded
2023, Paris, France
Founder / CEO
Gaël Delalleau (solo founder) — École Polytechnique; ex-offensive cybersecurity; 4× DEFCON CTF finalist; prior founder of Stribe (TC50 2009)
Funding
~$5M seed (secondary-sourced; TechCrunch disclosed no figure), co-led by Varsity VC
Other Backers
Scaleway, Bpifrance Deep Tech Program, French Tech 2030 label (Oct 2025)
Product
Kog Inference Engine (KIE): a low-level, single-kernel software layer for high single-request token throughput on standard GPUs; drop-in vLLM replacement
Team
11 people — ~10 engineers/researchers, 5 PhDs
Traction
Early access (“Request API Access”); ~200 business leads; unnamed “design partners” in app/game generation; no named paying customer
Valuation
Not disclosed

How “30x” Becomes “3.5x”

The headline number is engineered from the single most flattering configuration possible. Follow the qualifiers and the multiplier shrinks at every step.

The Path From Headline to Fine Print

01

“30x Faster”

The marketing headline on kog.ai. No model, batch size, or baseline attached to the number itself.

02

Batch Size 1

Measured at single-request decoding — the config that most flatters a memory-streaming engine and least reflects production economics.

03

A 2B Custom Model

The demo ran on Laneformer 2B, Kog’s own purpose-built coding model (~50% HumanEval) — not a standard frontier model.

04

A Theoretical Ceiling

3,000 tok/s on 8×MI300X — which Kog’s own blog labels an “upper bound, not guaranteed achievable,” at just 36% memory-bandwidth utilization.

05

“Up to 3.5x”

The number Kog’s own site cites for realistic AMD Instinct workloads. Its only real-model figure: 1,368 tok/s on Llama-3 8B.

06

Zero Third-Party Proof

No independent lab, no head-to-head vs. vLLM / TensorRT-LLM / SGLang, no named customer. Every number is self-reported.

To its credit, Kog’s engineering candor is real. The company open-sourced its demo model and its blog explicitly lists what was excluded from the benchmark — no quantization, no speculative decoding, no pruning, no KV-cache compression — and calls the figures upper bounds. The problem is not dishonesty in the fine print; it is a headline number chosen to be ~10x larger than the fine print the company itself publishes.

The Margin Kog Sells Belongs to Its Suppliers

Kog’s entire value proposition is defensible only as long as Nvidia (TensorRT-LLM), AMD (ROCm/vLLM), and open-source SGLang leave a single-kernel latency margin on the table. A hardware-specific inference engine is exactly the kind of optimization the chipmakers are motivated to absorb into their free, bundled software — margin that evaporates into the vendors’ stacks the moment they prioritize it. Kog has never publicly explained how an 11-person team sustains a moat against the very silicon vendors it depends on.

Squeezed From Both Sides

Kog is wedged between free first-party stacks below it and far-better-capitalized neutral rivals beside it — on roughly ~$5M of seed capital.

ZML

$20M seed

The closest analog — also French, also chip-agnostic (Nvidia/AMD/TPU/Metal/Intel), bypassing CUDA. Backed by LeCun, Hykes, and Hugging Face; named by TechCrunch as Kog’s direct rival, with 4x the capital.

Infinity

$15M seed @ ~$100M

“Software layer that makes any AI chip inference-ready,” auto-generating low-level kernels — the same wedge as Kog, with OpenAI/Anthropic-researcher backing and a stronger network.

Free

vLLM / TensorRT-LLM / SGLang / ROCm

The real threat: free, vendor-backed, improving every release cycle. The single-request latency Kog charges for is theirs to reclaim — and they are motivated to do so.

Scale

Fireworks $17.5B · Together $8.3B · Baseten ~$13B

The inference-cloud incumbents. They validate that inference is the 2026 capital magnet — and set the bar for how much a serious inference company is expected to raise.

HW

Groq $650M · Positron $230M · Cerebras

The “build new silicon” camp Kog defines itself against. Purpose-built inference hardware attacking the same latency problem from the metal up.

$5M

Kog’s Position

Barely enough to reach the September Series-A-gating demo, with no margin for error — a severe capital asymmetry against every player above.

Market timing is the strongest part of the thesis. “More performance from existing GPUs” is squarely in-thesis for the 2026 inference boom. But timing cuts both ways: the same boom has funded neutral rivals 3–4x deeper, and the incumbents whose silicon Kog optimizes are the ones best positioned to close the gap for free.

What the ~$5M Does Not Resolve

Seven structural concerns, each stated as a falsifiable claim.

High

Benchmark Cherry-Picking

The 30x rests on a batch-1, 2B custom-model, theoretical-ceiling comparison; the site itself shows “up to 3.5x.” Falsifiable: an independent Llama-3 8B run vs. optimized vLLM/TensorRT-LLM.

High

Moat Evaporates Into Vendor Stacks

A single-kernel, GPU-specific engine is exactly what Nvidia and AMD can absorb for free. Falsifiable: watch whether their next release ships comparable single-request latency.

High

Zero Independent Validation

All numbers self-reported; “200 leads” and unnamed “design partners” are not revenue. Falsifiable: a named, referenceable customer running a frontier model in production.

Medium

The Milestone Hasn’t Shipped

The entire Series A thesis hinges on a September “10x on a major model” demo that does not yet exist. Falsifiable: does it land, and at what memory-bandwidth utilization?

Medium

Batch-1 ≠ Production Economics

Real inference cost is throughput-at-concurrency; Kog publishes no cost-per-token or high-batch numbers. Falsifiable: concurrent-throughput and $/1M-token disclosures.

Medium

Undisclosed Dependencies

No ROCm/CUDA version, firmware, interconnect, or Scaleway-specific requirements published for the 8×GPU results. Falsifiable: reproducibility on a non-Scaleway MI300X node.

Medium

Sovereignty-Subsidy Concentration & Tiny Team

11 people, with capital and label tied to French/EU deep-tech policy; the thesis partly rests on compute-sovereignty politics rather than open-market performance wins. Falsifiable: commercial traction outside French/EU-subsidized channels.

Assessment Matrix

Technical Claim Credibility
Low-Medium
Real engineering and honest caveats, but 30x is a ceiling; no independent or head-to-head proof
Defensibility / Moat
Low
Sits in the crush zone between free vendor stacks and better-funded neutral rivals; no disclosed IP moat beyond kernel craft
Market Timing
High
Inference is the 2026 capital magnet; “more from existing GPUs” is squarely in-thesis
Team
Medium-High
Exceptional founder pedigree and PhD-dense, but solo founder + 11 people is thin for this fight
Funding Adequacy
Low
~$5M against $15M–$1.5B-funded rivals; barely enough to reach the Series-A-gating demo
Disclosure Quality
Medium
Honest technical blog, but no disclosed round size and a headline number ~10x its own fine print
Overall Investment Thesis
Low-Medium
High-variance deep-tech bet: credible talent and perfect timing, unverified claims and a structurally weak moat
Category
AI Inference Infra
Software-only LLM inference acceleration on standard datacenter GPUs

Kog is a real team, well-timed, taking the hardest way in. The engineering is credible and the candor in its technical blog is unusual. But the signature “30x” is an unverified batch-1 theoretical ceiling that Kog’s own website walks down to “3.5x,” the ~$5M seed is secondary-sourced and appears to be the same round as the October 2025 French Tech 2030 label — not a fresh raise — and the moat question is the one Kog has never answered: it sells the exact latency margin that Nvidia’s and AMD’s free first-party stacks are motivated to reclaim. The entire Series A thesis rests on a September milestone that has not shipped.

Research Sources

Based entirely on publicly available information, including the TechCrunch report of August 14, 2026. Every performance figure in this report is self-reported by Kog unless otherwise noted; the ~$5M round size is confirmed only via secondary aggregation, not a primary filing.

  1. TechCrunch — “Kog is going deeper to squeeze more inference out of GPUs” (Anna Heim, August 14, 2026) — investors, team size, competitor context; no dollar figure disclosed
  2. blog.kog.ai — “Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)” — benchmark figures, “upper bound” caveats, exclusion list
  3. blog.kog.ai — “Building a single-kernel, latency-optimized LLM inference engine on AMD MI300X GPUs” — technical architecture
  4. kog.ai — Company site: “30x,” “up to 3.5x on AMD Instinct,” “1,368 tokens/second on Llama-3 8B,” drop-in vLLM claim, early access
  5. Secondary aggregation (company/database summaries) — source of the ~$5M figure (Varsity VC + Bpifrance Deep Tech), 2023 founding, Paris HQ, October 2025 French Tech 2030 label
  6. PitchBook — Kog company profile (exists; no valuation surfaced)
  7. Competitor context: ZML ($20M seed, Dealroom/TechCrunch), Infinity ($15M seed @ ~$100M, BusinessWire), Together AI ($8.3B), Fireworks AI ($17.5B), Baseten (~$13B), Groq ($650M), Positron ($230M), Tenstorrent (~$1.2B)
  8. AMD ROCm / Nvidia TensorRT-LLM / vLLM / SGLang documentation — first-party inference-stack competitive analysis