Fish Audio, a Palo Alto-based AI voice startup, announced on July 29, 2026 that it has raised $52 million in a seed round led by Coreline Ventures and Capital Today, with additional participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
The company, founded by former NVIDIA researcher Shijia Liao and CEO Rissa Cao, offers a library of more than 15,000 natural language controls for AI voice generation. Since launching last year, Fish Audio has attracted more than 8 million users across its open-source and hosted platforms and now generates $21 million in annual recurring revenue. Its Fish Speech repository on GitHub has accumulated more than 31,000 stars.
Fish Audio has released five models over the past year — four speech generation models and one speech-to-text model. Three of the speech generation models are open-source, while its latest S2.1 Pro model is available only through a paid API. Enterprise customers including HeyGen and Sanas are already using the platform. The company offers monthly plans for creators and teams, as well as enterprise API access.
Cao said the company pursued outside capital to develop more advanced models and expand its enterprise offering as investor interest grew. Plans for the remainder of 2026 include releasing an audio understanding model and a speech-to-speech model.
The fundraise comes amid a notable trust issue for the platform. Earlier this year, some creators alleged their voices were uploaded to Fish Audio without their consent. The company has since automated its takedown process, with Cao stating that verified voice removal now takes less than three minutes. However, voices can still be uploaded without an artist’s knowledge until a removal request is filed.
Osuke Honda, a partner at lead investor Coreline Ventures, noted that the community-driven model depends on creator trust, calling for “verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models” to be built into the product.
Fish Audio competes in a crowded speech generation market that includes ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp.
Source: TechCrunch