AI Voice Startup Fish Audio Secures $52 Million in Seed Funding
  • News
  • North America

AI Voice Startup Fish Audio Secures $52 Million in Seed Funding

The funding will help the one-year-old startup scale its expressive, steerable AI voice models.

7/28/2026
Ghita Khalfaoui
Back to News

Fish Audio has raised $52 million in seed financing to expand its artificial intelligence voice platform for creators, developers, and enterprise customers. The round was led by Coreline Ventures and Capital Today, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners, and angel investors. Announced near the company’s first anniversary, the financing will support more advanced audio models and broader commercial operations.


From Open-Source Project to AI Voice Company

Fish Audio began as a research project developed by co-founder and chief scientist Shijia Liao, a former Nvidia researcher dissatisfied with the limited expressiveness of synthetic voices. Liao initially trained a voice-generation model on a single consumer-grade GPU before releasing the technology as an open-source project called Fish Speech. The project gained traction among developers, game designers, and creators, eventually becoming the foundation of a commercial platform led by CEO and co-founder Rissa Cao.

Rapid Growth in Its First Year

The company says it launched five audio models during its first year, including four text-to-speech systems and one speech-to-text model. Fish Audio also reports more than 8 million users, over 2 million community voice models, and annual recurring revenue of $21 million. Its workforce has grown from three people to 22, while enterprise and developer customers now account for roughly two-thirds of revenue.

Building More Expressive Voice Models

Fish Audio is positioning its technology as an alternative to text-to-speech systems that pronounce words accurately but often struggle with emotion, pacing, and natural interpretation. Its S2.1 Pro model offers more than 15,000 natural-language controls, supports over 83 languages, and allows users to adjust emotion at the word level. The company targets applications ranging from character voices and AI avatars to customer-service agents, sales tools, and real-time conversational products.

Expanding Enterprise Adoption

The startup offers paid plans for creators and teams, alongside enterprise APIs and deployment options for organizations with stricter performance, privacy, and security requirements. Fish Audio says its enterprise services include on-premises deployment, zero-data-retention policies, and configurations for regulated use cases, while companies such as HeyGen and Sanas are already using its technology. It is also deepening integrations with voice infrastructure companies including LiveKit and Retell to make expressive, low-latency speech available across more applications.

Funding Future Audio Research

Fish Audio plans to use the capital to develop an audio language model that can interpret sound and speech more broadly, as well as a speech-to-speech system. The company will also invest in developer tools, API infrastructure, its enterprise team, and partnerships as it competes with providers including ElevenLabs, Cartesia, Speechify, and WellSaid. Its ability to train and operate models efficiently could be important in a market where model quality, latency, control, and computing costs influence adoption.

Voice Ownership Remains a Challenge

The company’s community-driven voice library has also raised concerns after some creators alleged that their voices had been uploaded without permission. Fish Audio told TechCrunch that it has automated its takedown process, allowing people to submit a voice sample or contract and request removal, but the system still depends on affected individuals discovering unauthorized models. As the platform scales, stronger verification, licensing, attribution, and revenue-sharing mechanisms may become essential to maintaining creator trust.


The $52 million seed round gives Fish Audio substantial resources to move beyond its open-source origins and compete more aggressively in the voice AI market. Its early user growth, reported revenue, multilingual capabilities, and expanding enterprise business indicate demand for more expressive and controllable synthetic speech. The company’s next stage will depend on technical performance, commercial execution, and its ability to address consent and voice-ownership concerns as adoption broadens.