$20Starting price
0Popularity
Zro featured image

About Zro

Zro is a private inference endpoint designed specifically for open-weight coding models. It operates from EU-based infrastructure with a strict policy of zero request retention and prohibits training on customer data. The service supports popular coding agents such as Claude Code, Codex CLI, Cursor, Cline, and others through an npm package and APIs compatible with OpenAI or Anthropic standards. Zro is engineered for long-context, multi-turn coding sessions by leveraging HyperQuant compression and custom attention kernels optimized for AMD, NVIDIA, and TPU hardware. Currently, it offers access to models like MiniMax M3 and GLM-5.2, with plans to expand the roster of available open-weight models. Billing is structured with a starting point of $20 per month for $60 of inference spend, alongside pay-as-you-go usage packs for flexible consumption.

Key features

  • Private inference endpoint for open-weight coding models
  • Zero request retention and no training on customer data
  • EU-based infrastructure
  • Support for OpenAI- and Anthropic-compatible APIs
  • Integration with popular coding agents via npm package
  • HyperQuant compression for long-context sessions
  • Custom attention kernels optimized for AMD, NVIDIA, and TPU hardware
  • Access to MiniMax M3 and GLM-5.2 models
  • Pay-as-you-go usage packs available
  • Monthly billing starting at $20 for $60 of inference spend

Use cases

  • Running private AI coding model inference for secure development environments
  • Multi-turn coding sessions with long-context support
  • Integration with existing coding agents and workflows

Pros

  • Zero request retention and no training on customer data for privacy compliance
  • Optimized for long-context, multi-turn coding sessions with low latency
  • Supports multiple coding agents and tools via an npm package and OpenAI/Anthropic-compatible APIs
  • Multi-region infrastructure for improved performance and reliability
  • Leverages HyperQuant compression and custom attention kernels for enhanced speed

Cons

  • Limited to open-weight coding models, excluding proprietary alternatives
  • Requires setup with supported coding agents or manual configuration for some tools
  • Performance metrics may vary depending on model and region

Frequently asked questions about Zro

What is Zro?

Zro is a private inference endpoint for coding agents that serves open-weight models with zero request retention and no training on customer data.

Who can use Zro?

Developers and teams using coding agents like Claude Code, Codex CLI, Cursor, or Cline who require private, fast, and long-context inference.

How do I integrate Zro with my coding agent?

Install the @moonmath-ai/zro npm package, log in once, and launch supported tools with temporary session configurations using the CLI.

Does Zro support OpenAI or Anthropic-compatible clients?

Yes, Zro exposes OpenAI-compatible and Anthropic-compatible APIs for chat completions and messages requests.

What models does Zro currently support?

Zro supports models such as GLM-5.2, GLM-5.3 Flash, DeepSeek V4 Flash 0731, and Kimi K3, with plans to expand the roster.

Does Zro retain prompts or completions after inference?

No, Zro does not retain prompt or completion data after processing is complete.

Zro compared

Reviews