Fish Audio has secured $52 million in a seed round led by Coreline Ventures and Capital Today. The company aims to scale its library of 15,000 natural language controls, catering to diverse needs ranging from realistic AI avatars for HeyGen to expressive character voices for gaming studios.
Financial Traction and Funding
Since its launch last year, Fish Audio has scaled rapidly, attracting over 8 million users across its hosted and open-source offerings. This growth has translated into $21 million in annual recurring revenue. The recent $52 million investment round included participation from several firms, including Capital Today, 359 Capital, and HF0. While the company operated efficiently as an open-source project initially, CEO Rissa Cao noted that external capital is now necessary to develop advanced models and better support enterprise clients.
The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable.
Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup now has more than 8 million people using the open source or hosted versions of its models, and generates annual recurring revenue of $21 million.
To continue building on that traction, the startup on Tuesday said it has raised $52 million in a seed round that was led by Coreline Ventures and Capital Today.
Product Portfolio and Technology
Founded by former Nvidia researcher Shijia Liao, the company began as a single-GPU project focused on improving synthetic voice expressiveness. Over the past year, Fish Audio released five models: one for speech-to-text and four for speech generation. While three of these are open-source—with the Fish Speech repository earning over 31,000 GitHub stars—the S2.1 Pro model is restricted to a paid API. The company currently provides tiered monthly plans for creators and specialized APIs for enterprise partners like Sanas.
The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. Fish Audio started as a small project by former Nvidia researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice-generation model on a single GPU, which he then open-sourced.
It has open-sourced three of its speech-generation models, but its latest S2.1 Pro model is available only through its paid API.
Content Governance Challenges
The startup's method of compensating users for voice training data led to allegations that some voices were uploaded without consent. Although a DMCA process existed, slow response times prompted the company to automate takedowns. CEO Rissa Cao stated that creators can now prove ownership via contracts or voice samples to remove content in under three minutes. Despite this, Coreline Ventures partner Osuke Honda emphasized that long-term success requires industry-wide shifts toward verified voice ownership and transparent licensing.
Fish Audio offers paid monthly plans suited for creators and teams that unlock a set number of minutes of generation plus voice-cloning features. The company also offers an enterprise version of its APIs and platform, and says organizations like HeyGen and Sanas are already using it.
“Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls,” Fish Audio’s CEO and co-founder Rissa Cao said.
One way the startup has built its library of voices is by asking users to submit their own voices for training its models, and compensating them if their voices are used.
Strategic Roadmap
Fish Audio is positioning itself against competitors like ElevenLabs and Speechify by focusing on steerability and natural sound. To expand its capabilities, the company intends to launch an audio understanding model before the end of this year. Additionally, development is underway for a speech-to-speech model. These additions aim to meet specific enterprise demands, such as the low-latency requirements sought by voice agent firms like LiveKit.
That resulted in some trouble a few months ago, however, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA takedown process in place to address such concerns, but the takedowns themselves took a long time.
Key signals
- Rapid transition from open-source project to $21 million ARR.
- Automation of voice takedowns to mitigate consent disputes.
- Planned release of audio understanding and speech-to-speech models.
- Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls.
- “Every enterprise has different use cases and different preferences.
What to watch
Whether Fish Audio's automated removal process satisfies creator concerns regarding unauthorized uploads and if the new audio understanding model can differentiate them in a crowded market.
Source and methodology
This Intelligence Daily briefing preserves the key facts published by TechCrunch AI and organizes them into a fuller, reader-friendly report. Read the original reporting.