
Etched
Products · 2 tracked products
Sohu Transformer Inference ASIC
The world's first application-specific integrated circuit dedicated to Transformer architectures, embedding attention mechanisms directly into silicon. This chip achieves inference throughput exceeding 500,000 tokens per second on a single server, optimizing computational efficiency and latency for large language model deployment in data center environments.
Sohu Rack-Level Inference System
A rack-scale inference system co-designed around the Sohu chip, integrating hardware, packaging, interconnects, and software stacks. It targets large-scale deployment for trillion-parameter Mixture-of-Experts models and long-context scenarios, providing an optimized infrastructure solution for high-throughput AI serving in enterprise computing clusters.
Physix Frontier