Back to news
AI Tools & Products
Aug 14, 2026

Kog aims to enhance AI inference speed using existing GPUs through software optimization

Aug 14, 2026
AI Summary

French startup Kog is focusing on improving AI inference speeds on conventional GPUs, such as AMD and NVIDIA models, through software optimization. The company has garnered significant interest from businesses looking to reduce delays in AI workflows, with plans to demonstrate its technology on larger models in the near future.

  • Kog is a French startup that aims to optimize AI inference speed on standard GPUs, such as AMD MI300X and NVIDIA H200, rather than using purpose-built chips.
  • The company gained attention with a tech preview that showcased fast single-request decoding capabilities, attracting over 200 business leads.
  • Kog's CEO, Gaël Delalleau, believes that software engineering will be the initial use case for their technology, targeting customers who rely on AI for professional tasks and face delays.
  • The startup's Kog Inference Engine (KIE) aims to provide faster outcomes for users generating games and applications, potentially increasing revenue for these customers.
  • Kog is focusing on developing larger models to meet market demand, despite the challenges of delivering on its promise of 30x faster LLM inference.
  • The company has demonstrated impressive results with a small model, achieving 3,000 tokens per second, and is confident that similar results can be achieved with larger models.
  • Kog's approach involves deep-level GPU engineering research, which is time-consuming and limits the number of chips they can work with due to their small team size.
  • The startup is supported by Scaleway and backed by France’s Bpifrance and French Tech 2030 program, aiming to contribute to Europe's technological sovereignty.
  • Kog plans to demonstrate its technology on larger models by September to secure further funding and customer traction.
gpuinferenceagentic workflowsstartuptechnology