Google’s Frozen Chip: Gemini Baked into Silicon
A new report reveals that Google is developing a server chip with Gemini’s design directly integrated into the silicon. This innovative approach could significantly enhance efficiency and reshape the landscape of AI technology.
A New Paradigm for AI Chips
Currently, most AI chips are general-purpose, loading models onto them for execution. However, Google’s ambitious project, code-named "Frozen v2," aims to revolutionize this by creating a chip that embodies Gemini’s neural network architecture within its hardware. This means the chip’s structure remains fixed while engineers can update the model’s weights.
Efficiency and Performance Gains
The potential benefits are substantial. According to The Information, the Frozen v2 chip could be 6-10 times more efficient than Google’s current custom AI chips, measured by tokens served per unit of power. This efficiency gain is crucial as running AI models consumes vast amounts of energy, with every watt saved translating to significant cost savings at data center scale.
Strategic Implications
The timing of this project is significant. The Information suggests that Frozen v2 is a response to an AI capacity crunch within Google, leading to internal tensions and the rejection of some external customers. By creating a chip optimized for a specific model, Google can reduce overhead and improve performance, particularly in real-time applications like voice assistants where latency is critical.
Competition and Future Prospects
Google is not alone in exploring this technology. Startups like Taalas are already offering chips with directly printed model weights and architecture, such as their Hardcore chip. Taalas claims impressive performance, serving up to 17,000 tokens per second, a significant improvement over top Nvidia GPUs. This approach aims to address memory crunches and reduce the reliance on high-bandwidth memory.
If successfully implemented, Google’s Frozen chip, or similar technologies, could mark a significant shift in AI infrastructure, offering faster, more efficient, and more cost-effective solutions for running complex models at scale.