Back to news
Large Language Models
6d ago

DwarfStar 4 Launches Local Inference Engine for High-Memory Machines

Oct 2, 2026
AI Summary

DwarfStar 4 is a new inference engine designed for high-memory Mac, CUDA, and ROCm machines, supporting various models including DeepSeek V4 and GLM 5.x. It offers a local API, CLI, and a native agent, enabling users to run large models locally without relying on remote servers.

  • DwarfStar 4 (ds4) is a C inference engine optimized for high-memory environments, supporting models like DeepSeek V4, GLM 5.x, and Qwen3.8.
  • The engine operates in three phases: starting with a large mixture-of-experts model, followed by asymmetric quantization for efficiency, and culminating in a local engine with integrated APIs and agents.
  • Users can interact with the engine through a command line interface (CLI), HTTP APIs, and a native agent, all sharing the same model state and cache.
  • The system is designed to compress routed experts while maintaining critical paths, allowing for effective performance on specified hardware.
  • Users can download project GGUFs, build for their specific backend, and utilize the engine for local coding and model interactions.
  • The engine supports various execution modes based on platform and memory specifications, with benchmarks indicating strong performance for models like V4 Flash Q2 and GLM 5.3 Q2 at 128 GB memory.
llmredislocalds4open-source