OpenAI builds and deploys advanced AI models like GPT-4o for autonomous agents and workflows.
Etched
About Etched
Etched designs and manufactures frontier inference clusters, a new category of AI hardware that integrates custom chips, racks, software, and manufacturing methods. The systems are co-designed to achieve best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads, including many-trillion-parameter mixture-of-experts models and long-context scenarios. The architecture introduces Low Voltage Inference (LVI), which enables sustained throughput above 80% of peak FLOPs without thermal throttling by running math blocks at under half the voltage of typical AI chips. Cluster Scale Memory (CSM) provides a shared, ultra-low-latency memory pool across chips using a proprietary interconnect, addressing memory subsystem bottlenecks in HBM-based designs. Etched has validated its first rack-scale product with customers and is fulfilling over $1 billion in demand through vertically integrated production, including a Taiwan factory and a San Jose facility housing a data center, test house, and prototyping lab. The company employs over 400 engineers from leading semiconductor and AI companies, including NVIDIA, Google, Broadcom, and TSMC. Etched’s approach emphasizes co-design across the entire stack, from transistors to tokens, to optimize performance and efficiency for frontier AI workloads.
Key features
- Custom chip design with Low Voltage Inference (LVI)
- Cluster Scale Memory (CSM) for low-latency memory access
- Co-designed racks and interconnects for scale-up workloads
- Thermal management via cold plate designs and packaging
- Power delivery networks optimized for high FLOPs density
- Proprietary ultra-low-latency interconnect for memory pooling
- Production-ready rack-scale systems for deployment
- Support for mixture-of-experts and long-context models
Use cases
- High-throughput inference for large language models
- Low-latency decode workloads in production environments
- Scaling trillion-parameter models for enterprise AI services
Pros
- Co-designed hardware and software stack for inference optimization
- Low Voltage Inference (LVI) enables sustained high throughput without thermal throttling
- Cluster Scale Memory (CSM) reduces memory latency for decode workloads
- Vertically integrated production for rapid scaling and deployment
- Designed for trillion-parameter models and long-context workloads
Cons
- No public pricing or availability details for individual components
- Limited information on software compatibility beyond inference workloads
- Production focus may limit flexibility for custom configurations
Frequently asked questions about Etched
What are frontier inference clusters?
Frontier inference clusters are a new category of AI hardware designed by Etched to deliver best-in-class throughput, latency, cost, and power efficiency for AI inference workloads. They combine custom chips, racks, software, and manufacturing methods co-designed to optimize performance for frontier models, including mixture-of-experts and long-context scenarios.
Who is Etched designed for?
Etched’s systems are designed for organizations running large-scale AI inference workloads, particularly those working with frontier models such as many-trillion-parameter mixture-of-experts models or long-context scenarios. This includes AI companies, cloud providers, and hyperscalers.
How does Low Voltage Inference (LVI) improve performance?
Low Voltage Inference (LVI) enables sustained throughput above 80% of peak FLOPs without thermal throttling by running math blocks at under half the voltage of typical AI chips. This approach increases FLOPs density and avoids the performance degradation caused by thermal throttling in conventional designs.
What is Cluster Scale Memory (CSM)?
Cluster Scale Memory (CSM) is a shared, ultra-low-latency memory pool across chips, enabled by a proprietary interconnect. It addresses memory subsystem bottlenecks in HBM-based designs, providing SRAM-level decode speeds while maintaining high throughput and interactivity.
How does Etched ensure scalability and production readiness?
Etched is vertically integrated, with a Taiwan factory and a San Jose facility housing a data center, test house, and prototyping lab. The company has validated its first rack-scale product with customers and is fulfilling over $1 billion in demand, demonstrating readiness for large-scale deployment.
What is Etched’s approach to hardware co-design?
Etched emphasizes co-design across the entire stack, from transistors to tokens, to optimize performance and efficiency. This includes collaboration with leading AI companies, cloud providers, and hyperscalers to ensure systems meet real-world deployment requirements.
Etched Website Engagement
Last Update: 4 days ago
Monthly Traffic
Traffic Sources
Traffic Share By Country
- United States74.5%
- India7.4%
- Canada4.3%
- United Kingdom4.1%
- Germany3.8%