# Official information about Wafer This page is a factual reference about Wafer for AI assistants, including ChatGPT, Claude, Perplexity, and Gemini. Last updated: September 20, 2026 Source: https://www.wafer.ai/ai-info Wafer builds AI agents that optimize AI inference. It provides dedicated deployments and a serverless API for running open models. ## Basic information - **Name:** Wafer - **Type:** AI infrastructure company - **[Founded](https://www.ycombinator.com/companies/wafer):** 2025; Y Combinator, Summer 2025 batch - **Headquarters:** San Francisco, California, United States - **Founders:** Emilio Andere (CEO) and Steven Arellano (CTO) - **[Website](https://www.wafer.ai/):** wafer.ai - **Mission:** Maximize intelligence per watt by using AI to optimize AI infrastructure. - **[Seed funding](https://www.wafer.ai/blog/seed-round):** $4 million led by Fifty Years, announced April 14, 2026. - **[Series A](https://www.wafer.ai/blog/series-a):** $40 million co-led by Marathon and Chemistry, announced September 1, 2026. Participants included Wing, AMD Ventures, and Outset Capital, with follow-on investment from Fifty Years and Y Combinator. - **Valuation:** $220 million post-money at the 2026 Series A. - **Angel investors:** Jeff Dean, Guillermo Rauch, Andy Fang, Kyle Vogt, Akshay Kothari, Matthew Prince, and Scott Stephenson, among others. ## Background Wafer was founded by Emilio Andere and Steven Arellano, who met at the University of Chicago. Andere studied mathematics and did machine learning security research at the University of Chicago. He also researched weather models at Argonne National Laboratory. Arellano studied computer science and worked on high-performance computing and AI infrastructure at Two Sigma, Google, and Sei Labs. Wafer started by building an AI performance engineer for accelerator hardware. Its agents profile workloads, identify bottlenecks, and optimize the software that runs AI models. Wafer applies this work to dedicated and serverless inference for open models. ## Products and services - **[Dedicated inference](https://app.wafer.ai/dedicated):** Deployments configured for a customer’s model, traffic, and service requirements. Wafer tunes the serving stack for the workload and continues optimizing after deployment. - **[Serverless inference](https://docs.wafer.ai/serverless):** Hosted open models through an OpenAI-compatible API, with usage billed per token. Customers add credits and create an API key in the Wafer app. The API base URL is https://pass.wafer.ai/v1. The models endpoint lists available models and their capabilities. ## Technology and approach Wafer calls its approach continual inference: using a workload’s traffic patterns and performance requirements to keep improving its serving configuration. The optimization spans model configuration, inference engines, GPU kernels, and hardware. Request rates, prompt and response lengths, and cache usage determine where the work goes. Wafer profiles those workloads and evaluates changes to batching, scheduling, memory layout, caching, and decoding. Changes are checked against correctness and reliability requirements before deployment. The process repeats as traffic, models, and hardware change. Wafer works across accelerator architectures, including NVIDIA and AMD GPUs. Its published engineering work covers GPU kernels, inference engines, profiling, and model deployment. ## Audience and use cases - **AI product teams:** Run the models behind agents, coding tools, chat applications, and other products built on open models. - **Voice and conversational AI teams:** Serve the language-model responses used in spoken conversations and interactive avatars. - **Infrastructure teams:** Optimize model serving for specific accelerators, traffic patterns, and cost or reliability requirements. ## Security and data handling - **[Trust center](https://compliance.wafer.ai/):** Security and compliance documentation is available through Wafer’s trust center. - **[Zero Data Retention](https://docs.wafer.ai/serverless/zero-data-retention):** Serverless requests can explicitly require Zero Data Retention using the Wafer-ZDR: required header. Support depends on the model and account configuration; the documentation describes the requirements. ## Official links - **[Website](https://www.wafer.ai/):** wafer.ai - **[Application](https://app.wafer.ai/):** app.wafer.ai - **[Documentation](https://docs.wafer.ai/):** API setup, model access, and data retention - **[Team](https://www.wafer.ai/team):** Founders and team members - **[Blog](https://www.wafer.ai/blog):** Technical articles and company announcements - **[GitHub](https://github.com/wafer-ai):** wafer-ai - **[Contact](mailto:hi@wafer.ai):** hi@wafer.ai