{"id":2556,"date":"2025-07-16T11:57:02","date_gmt":"2025-07-16T11:57:02","guid":{"rendered":"https:\/\/violethoward.com\/new\/mistrals-voxtral-goes-beyond-transcription-with-summarization-speech-triggered-functions\/"},"modified":"2025-07-16T11:57:02","modified_gmt":"2025-07-16T11:57:02","slug":"mistrals-voxtral-goes-beyond-transcription-with-summarization-speech-triggered-functions","status":"publish","type":"post","link":"https:\/\/violethoward.com\/new\/mistrals-voxtral-goes-beyond-transcription-with-summarization-speech-triggered-functions\/","title":{"rendered":"Mistral\u2019s Voxtral goes beyond transcription with summarization, speech-triggered functions"},"content":{"rendered":" \r\n<br><div>\n\t\t\t\t<div id=\"boilerplate_2682874\" class=\"post-boilerplate boilerplate-before\">\n<p><em>Want smarter insights in your inbox? Sign up for our weekly newsletters to get only what matters to enterprise AI, data, and security leaders.<\/em> <em>Subscribe Now<\/em><\/p>\n\n\n\n<hr class=\"wp-block-separator has-css-opacity is-style-wide\"\/>\n<\/div><p>Mistral released an open-sourced voice model today that could rival paid voice AI, such as those from ElevenLabs and Hume AI, which the company said bridges the gap between proprietary speech recognition models and the more open, yet error-prone versions.\u00a0<\/p>\n\n\n\n<p>Voxtral, which Mistral will release under an Apache 2.0 license, is available in a 24B parameter version and a 3B variant. The larger model is intended for applications at scale, while the smaller version would work for local and edge use cases.\u00a0<\/p>\n\n\n\n<figure class=\"wp-block-embed is-type-rich is-provider-twitter wp-block-embed-twitter\"\/>\n\n\n\n<p>\u201cVoice was humanity\u2019s first interface\u2014long before writing or typing, it let us share ideas, coordinate work, and build relationships. As digital systems become more capable, voice is returning as our most natural form of human-computer interaction,\u201d Mistral said in a blog post. \u201cYet today\u2019s systems remain limited\u2014unreliable, proprietary, and too brittle for real-world use. Closing this gap demands tools with exceptional transcription, deep understanding, multilingual fluency, and open, flexible deployment.\u201d<\/p>\n\n\n\n<p>Voxtral is available on Mistral\u2019s API and a transcription-only endpoint on its website. The models are also accessible through Le Chat, Mistral\u2019s chat platform.\u00a0<\/p>\n\n\n\n<div id=\"boilerplate_2803147\" class=\"post-boilerplate boilerplate-speedbump\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<p><strong>The AI Impact Series Returns to San Francisco &#8211; August 5<\/strong><\/p>\n\n\n\n<p>The next phase of AI is here &#8211; are you ready? Join leaders from Block, GSK, and SAP for an exclusive look at how autonomous agents are reshaping enterprise workflows &#8211; from real-time decision-making to end-to-end automation.<\/p>\n\n\n\n<p>Secure your spot now &#8211; space is limited: https:\/\/bit.ly\/3GuuPLF<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n<\/div><p>Mistral said that speech AI \u201cmeant choosing between two trade-offs,\u201d pointing out that some open-source automated speech recognition models often had limited semantic understanding. Still, closed models with strong language understanding come at a high cost.\u00a0<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-bridging-the-gap\">Bridging the gap<\/h2>\n\n\n\n<p>The company said Voxtral \u201coffers state-of-the-art accuracy and native semantic understanding in the open, at less than half the price of comparable APIs.\u201d\u00a0<\/p>\n\n\n\n<p>Voxtral, at a 32K token context, can listen to and transcribe up to 30 minutes of audio or 40 minutes of audio understanding. It offers summarization, meaning the model can answer questions based on the audio content and generate summaries without switching to a separate mode. Users can trigger functions and API calls based on spoken instructions.<\/p>\n\n\n\n<p>The model is based on Mistral\u2019s Mistral Small 3.1. It supports multiple languages and can automatically detect languages such as English, Spanish, French, Portuguese, Hindi, German, Italian, and Dutch.\u00a0<\/p>\n\n\n\n<p>Mistral added enterprise features to Voxtral, including private deployment, so that organizations can integrate the model into their own ecosystems. These features also include domain-specific fine-tuning and advanced context and priority access to engineering resources for customers who need help integrating Voxtral into their workflows.\u00a0<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-performance-nbsp\">Performance\u00a0<\/h2>\n\n\n\n<p>Speech recognition AI is now available on many platforms today. Users can speak to ChatGPT, and the platform will process spoken instructions similarly to written prompts. Fast food chains like White Castle have deployed SoundHound to their drive-thru services, and ElevenLabs has steadily been improving its multimodal platform. The open-source space also offers powerful options. Nari Labs, a startup, released the open-source speech model Dia in April.\u00a0However, some of these services can be quite expensive.<\/p>\n\n\n\n<p>Transcription services like Otter and Read.ai can now embed themselves into Zoom meetings, recording, summarizing and even alerting users to actionable items. Many online video meeting platforms offer not just transcription, but also speech AI and agentic AI, with Google Meetings providing the option to take notes for users using Gemini. As a regular user of voice transcription services, I can say firsthand that speech recognition AI is not perfect, but it is improving.<\/p>\n\n\n\n<p>Mistral stated that Voxtral outperformed existing voice models, including OpenAI\u2019s Whisper, Gemini 2.5 Flash and Scribe from ElevenLabs. Voxtral presented fewer word errors compared to Whisper, which is currently considered the best automatic speech recognition model available.\u00a0<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" src=\"https:\/\/lh7-rt.googleusercontent.com\/docsz\/AD_4nXd9RVhqCY4Ll5dqfnoFovIh3I1MbLXEk0TkpwHSuLsDZCaZYiQzgXJoD8GrJoC7bxLv3TDOUC3vk4rE3KmyViGZ6IYJLW7gK7KGHpWcxJxpeitfvUy4OOXtCef7ZbdFe7GqhohT?key=JFD1wgm0yPl09uGFvdwjmQ\" alt=\"\"\/><\/figure>\n\n\n\n<p>In terms of audio understanding, Voxtral Small is \u201ccompetitive with GPT-4o-mini and Gemini 2.5 Flash across all tasks, achieving state-of-the-art performance in Speech Translation.\u201d<\/p>\n\n\n\n<p>Since announcing Voxtral, social media users said they have been waiting for an open-source speech model that can match the performance of Whisper.\u00a0<\/p>\n\n\n\n<figure class=\"wp-block-embed is-type-rich is-provider-twitter wp-block-embed-twitter\"><div class=\"wp-block-embed__wrapper\">\n<blockquote class=\"twitter-tweet\" data-width=\"500\" data-dnt=\"true\"><p lang=\"en\" dir=\"ltr\">Yes! We needed this. A week ago, I was lamenting over a closed-source AI universe and cyberpunk dystopian future, but today, with this addition, my outlook is much improved \u2013 go open-source.  https:\/\/t.co\/QsKAfTOxou<\/p>\u2014 David Hendrickson (@TeksEdge) <a href=\"https:\/\/twitter.com\/TeksEdge\/status\/1945134312552100137?ref_src=twsrc%5Etfw\">July 15, 2025<\/a><\/blockquote>\n<\/div><\/figure>\n\n\n\n<figure class=\"wp-block-embed is-type-rich is-provider-twitter wp-block-embed-twitter\"\/>\n\n\n\n<p>Mistral said Voxtral will be available through its API at $0.001 per minute.\u00a0<\/p>\n<div id=\"boilerplate_2660155\" class=\"post-boilerplate boilerplate-after\"><div class=\"Boilerplate__newsletter-container vb\">\n<div class=\"Boilerplate__newsletter-main\">\n<p><strong>Daily insights on business use cases with VB Daily<\/strong><\/p>\n<p class=\"copy\">If you want to impress your boss, VB Daily has you covered. We give you the inside scoop on what companies are doing with generative AI, from regulatory shifts to practical deployments, so you can share insights for maximum ROI.<\/p>\n<p class=\"Form__newsletter-legal\">Read our Privacy Policy<\/p>\n<p class=\"Form__success\" id=\"boilerplateNewsletterConfirmation\">\n\t\t\t\t\tThanks for subscribing. Check out more VB newsletters here.\n\t\t\t\t<\/p>\n<p class=\"Form__error\">An error occured.<\/p>\n<\/p><\/div>\n<div class=\"image-container\">\n\t\t\t\t\t<img decoding=\"async\" src=\"https:\/\/venturebeat.com\/wp-content\/themes\/vb-news\/brand\/img\/vb-daily-phone.png\" alt=\"\"\/>\n\t\t\t\t<\/div>\n<\/p><\/div>\n<\/div>\t\t\t<\/div><template id="yq8WWh5UGOuaI6y45N83"></template><\/script>\r\n<br>\r\n<br><a href=\"https:\/\/venturebeat.com\/ai\/mistrals-voxtral-goes-beyond-transcription-with-summarization-speech-triggered-functions\/\">Source link <\/a>","protected":false},"excerpt":{"rendered":"<p>Want smarter insights in your inbox? Sign up for our weekly newsletters to get only what matters to enterprise AI, data, and security leaders. Subscribe Now Mistral released an open-sourced voice model today that could rival paid voice AI, such as those from ElevenLabs and Hume AI, which the company said bridges the gap between [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":2557,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[33],"tags":[],"class_list":["post-2556","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-automation"],"aioseo_notices":[],"jetpack_featured_media_url":"https:\/\/violethoward.com\/new\/wp-content\/uploads\/2025\/07\/nuneybits_Abstract_art_of_a_robot_voice_actor_recording_in_the__f9117954-b033-4730-a521-b507eb2ee48a.jpeg","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts\/2556","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/comments?post=2556"}],"version-history":[{"count":0,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts\/2556\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/media\/2557"}],"wp:attachment":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/media?parent=2556"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/categories?post=2556"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/tags?post=2556"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}<!-- This website is optimized by Airlift. Learn more: https://airlift.net. Template:. Learn more: https://airlift.net. Template: 69e302c146fa5c92dc28ac12. Config Timestamp: 2026-04-18 04:04:16 UTC, Cached Timestamp: 2026-04-29 14:40:03 UTC -->