Skip to main content
6h ago

Fish Audio Nabs $50M Seed to Power AI Voice for Creators & Businesses

The burgeoning market for AI-generated voice models is experiencing significant growth. Creative applications increasingly demand AI voices that are h

4 min read10 views5 tags
Originally reported bytechcrunch

The burgeoning market for AI-generated voice models is experiencing significant growth. Creative applications increasingly demand AI voices that are highly expressive, while enterprises aiming to streamline customer support and sales operations require models that offer greater steerability and control.

Palo Alto-based Fish Audio is positioning itself to address this broad spectrum of needs with its extensive library featuring over 15,000 natural language controls. Since its launch last year, the startup has rapidly gained traction, with more than 8 million individuals now utilizing either the open-source or hosted versions of its models. This success is reflected in its current annual recurring revenue, which stands at $21 million.

To further capitalize on this momentum, the company announced on Tuesday that it has successfully secured $50 million in a seed funding round. The investment was co-led by Coreline Ventures and Capital Today, with additional participation from a diverse group of investors including 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.

Fish Audio originated as a passion project by former NVIDIA researcher Shijia Liao. Frustrated by the lack of expressiveness in existing synthetic voices, Liao independently trained a voice generation model on a single GPU and subsequently open-sourced it. This "Fish Speech" repository on GitHub has since garnered significant attention, boasting over 31,000 stars and becoming a valuable resource for indie developers, video game designers, and content creators.

Over the past year, the company has introduced five distinct models: four dedicated to speech generation and one for speech-to-text conversion. While three of its speech generation models remain open-source, its most advanced offering, the S2.1 Pro model, is exclusively accessible through its paid API.

Fish Audio provides subscription-based monthly plans tailored for individual creators and teams, offering a specified number of generation minutes along with advanced voice cloning capabilities. Furthermore, the company extends an enterprise version of its APIs and platform, which is already being utilized by prominent organizations such as HeyGen, Sanas, and Plaud.

“Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls,” stated Rissa Cao.

Historically, the startup expanded its voice library by inviting users to submit their own voices for model training, offering compensation if their submissions were utilized. However, this approach led to issues a few months ago when some creators alleged their voices were uploaded to Fish Audio without proper consent. Although the startup had a DMCA content takedown process in place to address these concerns, the resolution of takedown requests proved to be time-consuming.

In response to these challenges, Fish Audio’s CEO and co-founder, Rissa Cao, informed TechCrunch that the company has now fully automated its takedown process. Creators can easily provide a brief voice sample or a contract to verify ownership of an uploaded voice, ensuring its removal from the platform in under 3 minutes, she affirmed.

Nonetheless, this enhanced process does not inherently prevent the initial unauthorized upload of an artist's voice. Until the artist becomes aware and formally requests its removal, their voice could potentially remain in use on the platform.

Oskue Honda, a partner at Coreline Ventures, emphasized that the efficacy of a community-driven model is fundamentally contingent upon the trust creators place in the platform.

“A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially,” he articulated.

Cao explained that when the startup primarily offered its product as an open-source project with creator-focused plans, it operated efficiently without requiring external funding. However, the ambition to develop more advanced models and cater to enterprise clients, coupled with increasing investor interest, prompted the company to seek capital.

Looking ahead, Fish Audio has outlined plans to release an audio understanding model later this year. Additionally, the company is actively developing a sophisticated speech-to-speech model.

The speech generation market is highly competitive, featuring numerous players such as ElevenLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle), and Krisp, all vying for the business of creators and enterprises alike.

Rico Mallozzi, a partner at 359 Capital, believes that Fish Audio's provision of fine-grained controls for developers and its ability to train models cost-efficiently will enable it to effectively compete with larger AI laboratories.

“I think what they’ve been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices,” Mallozzi conveyed to TechCrunch during a call.

#AI News#Fish Audio#AI Voice#Seed Funding#Open Source
ES
Editorial StaffEditor

The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.

View all posts
Reader feedback

What did you think of this story?

User Comments

Filter:
No comments yet. Be the first to comment!
Continue reading
View all news