A high-throughput and memory-efficient inference and serving engine for LLMs
-
Updated
Sep 9, 2026 - Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Never stop coding. Free MIT AI gateway: one endpoint, 352 providers (150+ free), 1200+ models Kimi, Claude, GPT, Gemini, GLM, DeepSeek, MiniMax. Works with Claude Code, Codex, Cursor, OpenCode, Cline & Copilot. Quota-aware auto-fallback, RTK+Caveman compression saves 15-95% tokens, MCP/A2A, Desktop/PWA. Built by 550+ contributors
The unified workspace where open-source models get things done for you.
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
✨ AI Coding, Vim Style
Command Code AI — the best coding agent for open models.
Visible multi-agent CLI workspace for mixing Codex, Claude, Gemini, Kimi, Qwen, Cursor, Copilot, Pi, OpenCode, and other AI coding agents
Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace viewer.
External-model router for Codex with guided Kimi OAuth/API, DeepSeek, safe migration, and rollback.
Find, benchmark and install in CLI 170+ FREE coding LLM models across 15+ providers in real time
Delegate a coding task to a separate coding agent CLI, review the diff, land the commit yourself — one per implementer.
U-Claw 虾盘 — OpenClaw AI 助手离线安装 U 盘:一键装好 Claude Code/Codex/OpenClaw,无需联网即插即用;新品 U-King AI 装机管家详见 u-king.org,另有远程支持与定制 AI 开发。
Web, Desktop & Mobile client for Codex, Claude Code, OpenCode, Kimi, Augment Code, Qwen, fully end-to-end encrypted
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
A macOS menu bar application that monitors AI coding assistant usage quotas. Keep track of your Claude, Codex, Antigravity ,and Gemini usage at a glance.
AI-native agent harness for coding workflows by python: multi-model LLM orchestration, stateful sessions, tool governance, traceable delivery, and provider routing for GPT, Claude, DeepSeek, Qwen, Kimi, GLM, and MiniMax.
Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere
To associate your repository with the kimi topic, visit your repo's landing page and select "manage topics."