- Deepslate has raised €7.7 million in seed funding, with 42CAP as the lead and Alstin Capital also participating.
- The laboratory, based in Berlin, has developed its own speech-to-speech model and runs it on servers in Germany.
- The response time of Deepslate was measured by Artificial Analysis at 0.44 seconds, placing it just behind Krafton’s Raon.
The voice AI lab Deepslate announced it had raised a seed round of €7.7 million, with 42CAP from Munich as the lead investor. Alstin Capital, SIVentures, and several business angels also participated in the round.
Prior to founding the Berlin-based startup in 2024, both co-founders, Paskal Paesler, the chief executive officer, and Jan Brachthäuser, the chief technology officer, began writing computer programs when they were 11 years old.
As Paesler shares with Tech Funding News, each of them had developed and led one of the two largest Minecraft servers in Europe. It was through this experience that Paesler became interested in AI, and later, he led a team of over 10 people at Stromee.
Paesler states that the majority of voice AI, including ElevenLabs and about 90% to 95% of the market, uses three stages: speech-to-text, language modelling, and text-to-speech. He points out that each of these stages introduces error and results in the loss “of all the sentiment, the way it is pronounced, generally speaking, the prosody, and the timbre of the voice.”
Two years ago, the founders came across projection models — something that, according to Paesler, larger labs had neglected — and changed their focus. Currently, Deepslate has 17 employees.
A fast model, but with some limitations
As Paesler explains, Deepslate has developed models that can understand and generate speech directly, without translating it into text. The model includes a speech encoder, a reasoning core based on an open-weights language model that has been post-trained for specific languages, and a speech decoder.
There are trainable projectors that link these parts together. If a better base model becomes available, Deepslate only retrains the projectors, reducing retraining time from months to days. The base model has not been disclosed.
Deepslate states that it takes 250 milliseconds from the end of a user’s sentence to produce a reply. The independent benchmarking company Artificial Analysis found a time to first audio of 0.44 seconds. On the Big Bench Audio, a reasoning test presented as speech, Opal, Deepslate’s model, achieves a score of about 85%. The claim that it has the best CoVoST2 error rate is based on the methodology published by Deepslate itself.
The announcement states that Opal is the fastest speech-to-speech model in the world.
Paesler regards OpenAI, Google, and Grok as his main competitors since these companies have launched speech-to-speech models. ElevenLabs raised $500 million at an $11 billion valuation in February 2026 and could be sold for around $22 billion. In Paris, Gradium raised $100 million in its seed round in July 2026, approximately 11 times Deepslate’s.
Deepgram completed a $130 million Series C, and Berlin’s Synthflow raised a $20 million Series A with Accel as the lead investor. What sets Deepslate apart is its use of a single model to process audio directly.
Sovereignty is the main selling point
Deepslate runs its models in Deutsche Telekom’s cloud in Germany and, if required, customers can also deploy the model on their own infrastructure, including in air-gapped environments. The company has ISO 27001 certification, which is significant for insurers, banks, and public authorities that often cannot use US-based cloud services.
This is in keeping with TFN‘s emphasis on sovereign startups, and voice startups are also performing better than the big tech firms in listening tests.
“Even before Deepslate, Paskal and Jan scaled latency-critical systems for millions of users. Today they are applying that experience to their own speech-to-speech model, which enables natural real-time conversations that are secure and run locally on the customer’s infrastructure. We invested because we see a globally competitive voice AI model lab that will build a central European solution for voice AI,” says Andreas Schenk, partner at Alstin Capital.
In December 2024, Alstin Capital completed raising its third fund at a value of €175 million, co-led the €10 million Series A financing of VoiceLine, a company based in Munich, and led the $20 million seed round for NeuralTrust, a firm based in Barcelona.
“Europe needs its own voice models, and Deepslate is one of the few teams here training its
own speech-to-speech model. What convinced us about Paskal and Jan is their technical depth: with a small team and a fraction of the compute used by large labs, they have built the fastest speech-to-speech model in the world,” adds Julian von Fischer, general partner at 42CAP.
The fund has recently acted as the lead investor in the $4 million seed round for Kontext, a company based in Munich, and in the €4 million seed round for ARC Intelligence, which is based in Berlin, and also, in association with Mozilla Ventures, provided the $3.2 million seed investment to Galtea, a company based in Barcelona.
The company states that insurers, contact centres, and various platforms are already using the model in production, even though it doesn’t mention any specific customers. Paesler estimates that the total size of the voice market is “probably” one trillion dollars, although this figure is not backed by a source. He also notes that Deepslate is the only European model in the benchmark he shared, out of six companies globally.
Deepslate claims that a small team can remain competitive by retraining thin layers on top of the best available open models. But how long can a small lab maintain its speed advantage as OpenAI, Google, and xAI continue to release real-time models, and will European buyers pay for verified hosting once that advantage diminishes?