{"id":3418,"date":"2025-08-29T05:47:50","date_gmt":"2025-08-29T05:47:50","guid":{"rendered":"https:\/\/violethoward.com\/new\/nous-research-drops-hermes-4-ai-models-that-outperform-chatgpt-without-content-restrictions\/"},"modified":"2025-08-29T05:47:50","modified_gmt":"2025-08-29T05:47:50","slug":"nous-research-drops-hermes-4-ai-models-that-outperform-chatgpt-without-content-restrictions","status":"publish","type":"post","link":"https:\/\/violethoward.com\/new\/nous-research-drops-hermes-4-ai-models-that-outperform-chatgpt-without-content-restrictions\/","title":{"rendered":"Nous Research drops Hermes 4 AI models that outperform ChatGPT without content restrictions"},"content":{"rendered":" \r\n<br><div>\n\t\t\t\t<div id=\"boilerplate_2682874\" class=\"post-boilerplate boilerplate-before\">\n<p><em>Want smarter insights in your inbox? Sign up for our weekly newsletters to get only what matters to enterprise AI, data, and security leaders.<\/em> <em>Subscribe Now<\/em><\/p>\n\n\n\n<hr class=\"wp-block-separator has-css-opacity is-style-wide\"\/>\n<\/div><p>Nous Research, a secretive artificial intelligence startup that has emerged as a leading voice in the open-source AI movement, quietly released Hermes 4 on Monday, a family of large language models that the company claims can match the performance of leading proprietary systems while offering unprecedented user control and minimal content restrictions.<\/p>\n\n\n\n<p>The release represents a significant escalation in the battle between open-source AI advocates and major technology companies over who should control access to advanced artificial intelligence capabilities. Unlike models from OpenAI, Google, or Anthropic, Hermes 4 is designed to respond to nearly any request without the safety guardrails that have become standard in commercial AI systems.<\/p>\n\n\n\n<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">Nous Research presents Hermes 4, our latest line of hybrid reasoning models.https:\/\/t.co\/E5EW9hBurb<\/p><p>Hermes 4 builds on our legacy of user-aligned models with expanded test-time compute capabilities. <\/p><p>Special attention was given to making the models creative and interesting to\u2026 <a href=\"https:\/\/t.co\/52VjnvrDWM\">pic.twitter.com\/52VjnvrDWM<\/a><\/p>\u2014 Nous Research (@NousResearch) <a href=\"https:\/\/twitter.com\/NousResearch\/status\/1960416954457710982?ref_src=twsrc%5Etfw\">August 26, 2025<\/a><\/blockquote> \n\n\n\n<p>\u201cHermes 4 builds on our legacy of user-aligned models with expanded test-time compute capabilities,\u201d Nous Research announced on X (formerly Twitter). \u201cSpecial attention was given to making the models creative and interesting to interact with, unencumbered by censorship, and neutrally aligned while maintaining state of the art level math, coding, and reasoning performance for open weight models.\u201d<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-how-hermes-4-s-hybrid-reasoning-mode-outperforms-chatgpt-and-claude-on-math-benchmarks\">How Hermes 4\u2019s \u2018hybrid reasoning\u2019 mode outperforms ChatGPT and Claude on math benchmarks<\/h2>\n\n\n\n<p>Hermes 4 introduces what Nous Research calls \u201chybrid reasoning,\u201d allowing users to toggle between fast responses and deeper, step-by-step thinking processes. When activated, the models generate their internal reasoning within special <code><think\/><\/code> tags before providing a final answer \u2014 similar to OpenAI\u2019s o1 reasoning models but with full transparency into the AI\u2019s thought process.<\/p>\n\n\n\n<div id=\"boilerplate_2803147\" class=\"post-boilerplate boilerplate-speedbump\">\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<p><strong\/><strong>AI Scaling Hits Its Limits<\/strong><\/p>\n\n\n\n<p>Power caps, rising token costs, and inference delays are reshaping enterprise AI. Join our exclusive salon to discover how top teams are:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Turning energy into a strategic advantage<\/li>\n\n\n\n<li>Architecting efficient inference for real throughput gains<\/li>\n\n\n\n<li>Unlocking competitive ROI with sustainable AI systems<\/li>\n<\/ul>\n\n\n\n<p><strong>Secure your spot to stay ahead<\/strong>: https:\/\/bit.ly\/4mwGngO<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n<\/div><p>The technical achievement is substantial. In testing, Hermes 4\u2019s largest 405-billion parameter model scored 96.3% on the MATH-500 benchmark in reasoning mode and 81.9% on the challenging AIME\u201924 mathematics competition \u2014 performance that rivals or exceeds many proprietary systems costing millions more to develop.<\/p>\n\n\n\n<p>\u201cThe challenge is making thinking traces useful and verifiable without runaway reasoning,\u201d noted AI researcher Rohan Paul on X, highlighting one of the technical breakthroughs in the release.<\/p>\n\n\n\n<p>Perhaps most notably, Hermes 4 achieved the highest score among all tested models on \u201cRefusalBench,\u201d a new benchmark Nous Research created to measure how often AI systems refuse to answer questions. The model scored 57.1% in reasoning mode, significantly outperforming GPT-4o (17.67%) and Claude Sonnet 4 (17%).<\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img fetchpriority=\"high\" decoding=\"async\" width=\"572\" height=\"562\" src=\"https:\/\/venturebeat.com\/wp-content\/uploads\/2025\/08\/GzS-zJWa4AEonw_.png\" alt=\"\" class=\"wp-image-3016215\" srcset=\"https:\/\/venturebeat.com\/wp-content\/uploads\/2025\/08\/GzS-zJWa4AEonw_.png 572w, https:\/\/venturebeat.com\/wp-content\/uploads\/2025\/08\/GzS-zJWa4AEonw_.png?resize=300,295 300w, https:\/\/venturebeat.com\/wp-content\/uploads\/2025\/08\/GzS-zJWa4AEonw_.png?resize=52,52 52w, https:\/\/venturebeat.com\/wp-content\/uploads\/2025\/08\/GzS-zJWa4AEonw_.png?resize=400,393 400w\" sizes=\"(max-width: 572px) 100vw, 572px\"\/><figcaption class=\"wp-element-caption\">Hermes 4 models from Nous Research answered significantly more questions than competing AI systems on RefusalBench, a test measuring how often models refuse to respond to user requests. (Credit: Nous Research)<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-inside-dataforge-and-atropos-the-breakthrough-training-systems-behind-hermes-4-s-capabilities\">Inside DataForge and Atropos: The breakthrough training systems behind Hermes 4\u2019s capabilities<\/h2>\n\n\n\n<p>Behind Hermes 4\u2019s capabilities lies a sophisticated training infrastructure that Nous Research has developed over several years. The models were trained using two novel systems: DataForge, a graph-based synthetic data generator, and Atropos, an open-source reinforcement learning framework.<\/p>\n\n\n\n<p>DataForge creates training data through what the company describes as \u201crandom walks\u201d through directed graphs, transforming simple pre-training data into complex instruction-following examples. The system can, for instance, take a Wikipedia article and transform it into a rap song, then generate questions and answers based on that transformation.<\/p>\n\n\n\n<p>Atropos, meanwhile, operates like hundreds of specialized training environments where AI models practice specific skills\u2014mathematics, coding, tool use, and creative writing\u2014receiving feedback only when they produce correct solutions. This \u201crejection sampling\u201d approach ensures that only verified, high-quality responses make it into the training data.<\/p>\n\n\n\n<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">Atropos is Nous&#8217; Reinforcement Learning framework<\/p><p>Atropos is an open source reinforcement learning environment by Nous that has hundreds of \u201cgyms\u201d (like math, coding, games, tool\u2011use, vision) to train and evaluate LLM trajectories via scalable, async RL loops.<\/p><p>In other words\u2026 <a href=\"https:\/\/t.co\/fjxaQKClEZ\">pic.twitter.com\/fjxaQKClEZ<\/a><\/p>\u2014 Tommy (@Shaughnessy119) <a href=\"https:\/\/twitter.com\/Shaughnessy119\/status\/1960477672007721223?ref_src=twsrc%5Etfw\">August 26, 2025<\/a><\/blockquote> \n\n\n\n<p>\u201cNous used these environments to generate the dataset for Hermes 4!\u201d explained Tommy Shaughnessy, a venture capitalist at Delphi Ventures who has invested in Nous Research. \u201cAll in the dataset contains 3.5 million reasoning samples and 1.6 million non-reasoning samples! Hermes was trained on RL data, not just static datasets of question and answer!\u201d<\/p>\n\n\n\n<p>The training process required 192 Nvidia B200 GPUs and 71,616 GPU hours for the largest model \u2014 a significant but not unprecedented computational investment that demonstrates how specialized techniques can compete with the massive scale of tech giants.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-why-nous-research-believes-ai-safety-guardrails-are-annoying-as-hell-and-hurt-innovation\">Why Nous Research believes AI safety guardrails are \u2018annoying as hell\u2019 and hurt innovation<\/h2>\n\n\n\n<p>Nous Research has built its reputation on a philosophy that puts user control above corporate content policies. The company\u2019s models are designed to be \u201csteerable,\u201d meaning they can be fine-tuned or prompted to behave in specific ways without the rigid safety constraints that characterize commercial AI systems.<\/p>\n\n\n\n<p>\u201cHermes 4 is not shackled by disclaimers, rules and being overly cautious which is annoying as hell and hurts innovation and usability,\u201d wrote Shaughnessy in a detailed thread analyzing the release. \u201cIf its open source but refuses all requests its pointless. Not an issue with Hermes 4.\u201d<\/p>\n\n\n\n<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">Hermes 4 is not shackled by disclaimers, rules and being overly cautious which is annoying as hell and hurts innovation and usability.<\/p><p>Hermes 4 70B is at the complete opposite of the spectrum vs OpenAI&#8217;s open source model. It&#8217;s also ~4x more open vs ChatGPT 4o!<\/p><p>If its open\u2026 <a href=\"https:\/\/t.co\/q5RpX1oOzo\">pic.twitter.com\/q5RpX1oOzo<\/a><\/p>\u2014 Tommy (@Shaughnessy119) <a href=\"https:\/\/twitter.com\/Shaughnessy119\/status\/1960477656388132974?ref_src=twsrc%5Etfw\">August 26, 2025<\/a><\/blockquote> \n\n\n\n<p>This approach has made Nous Research popular among AI researchers and developers who want maximum flexibility, but it also places the company at the center of ongoing debates about AI safety and content moderation. While the models can theoretically be used for harmful purposes, Nous Research argues that transparency and user control are preferable to corporate gatekeeping.<\/p>\n\n\n\n<p>The company\u2019s technical report, released alongside the models, provides unprecedented detail about the training process, evaluation results, and even the actual text outputs from benchmark tests. \u201cWe believe this report sets a new standard for transparency in benchmarking,\u201d the company stated.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-how-a-small-startup-with-192-gpus-is-competing-against-big-tech-s-billion-dollar-ai-budgets\">How a small startup with 192 GPUs is competing against Big Tech\u2019s billion-dollar AI budgets<\/h2>\n\n\n\n<p>Hermes 4\u2018s release comes at a pivotal moment in the AI industry. While major technology companies have poured billions into developing increasingly powerful AI systems, a growing open-source movement argues that these capabilities should not be controlled by a handful of corporations.<\/p>\n\n\n\n<p>Recent months have seen significant advances in open-source AI, with models like Meta\u2019s Llama 3.1, DeepSeek\u2019s R1, and Alibaba\u2019s Qwen series achieving performance that rivals proprietary systems. Hermes 4 represents another step in this progression, particularly in the area of reasoning\u2014long considered a strength of closed systems like OpenAI\u2019s o1.<\/p>\n\n\n\n<p>\u201cFirst up, Nous is a startup with dozens of extremely talented people,\u201d noted Shaughnessy. \u201cThey do not have the $100b+ annual capex spend of a hyperscaler nor 1,000\u2019s of employees and despite that they continue to put out innovative models and research at an insane pace.\u201d<\/p>\n\n\n\n<p>The startup, which raised $65 million in funding earlier this year led by Paradigm, has also been developing Psyche Network, a distributed training system that aims to coordinate AI training across internet-connected computers using blockchain technology.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-the-technical-fix-that-stopped-hermes-4-from-thinking-in-endless-loops\">The technical fix that stopped Hermes 4 from thinking in endless loops<\/h2>\n\n\n\n<p>One of Hermes 4\u2018s most significant technical contributions addresses a problem plaguing reasoning models: overly long thinking processes. The researchers found that their smaller 14-billion parameter model would reach maximum context length 60% of the time when reasoning, essentially getting stuck in endless loops of thinking.<\/p>\n\n\n\n<p>Their solution involved a second training stage that teaches models to stop reasoning at exactly 30,000 tokens, reducing overlong generation by 65-79% while maintaining most of the reasoning performance. This \u201clength control\u201d technique could prove valuable for the broader AI research community.<\/p>\n\n\n\n<p>\u201cSmaller models (&lt;14B) tend to overthink when distilled, but larger models don\u2019t,\u201d observed AI researcher Muyu He on X, highlighting insights from the technical report.<\/p>\n\n\n\n<p>However, Hermes 4 still faces limitations common to open-source models. Despite impressive benchmark performance, the models require significant computational resources to run and may not match the ease of use or reliability of commercial AI services for many applications.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-where-to-try-hermes-4-and-what-it-costs-compared-to-chatgpt-and-claude\">Where to try Hermes 4 and what it costs compared to ChatGPT and Claude<\/h2>\n\n\n\n<p>Nous Research has made Hermes 4 available through multiple channels, reflecting the open-source philosophy. The model weights are freely downloadable on Hugging Face, while the company also offers API access through its revamped chat interface and partnerships with inference providers like Chutes, Nebius, and Luminal.<\/p>\n\n\n\n<p>\u201cYou can try Hermes 4 in the new, revamped Nous Chat UI,\u201d the company announced, highlighting features like parallel interactions and a memory system.<\/p>\n\n\n\n<p>For enterprise users and researchers, the models represent a potentially attractive alternative to paying for API access to proprietary systems, especially for applications requiring high levels of customization or handling of sensitive content.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-the-bigger-picture-what-hermes-4-means-for-the-future-of-ai-development\">The bigger picture: What Hermes 4 means for the future of AI development<\/h2>\n\n\n\n<p>The release of Hermes 4 represents more than just another AI model launch \u2014 it\u2019s a statement about who should control the future of artificial intelligence. In an industry increasingly dominated by a handful of tech giants with virtually unlimited resources, Nous Research has demonstrated that innovation can still come from unexpected places.<\/p>\n\n\n\n<p>The company\u2019s approach raises fundamental questions about the trade-offs between safety and capability, between corporate control and user freedom. While major technology companies argue that careful content moderation and safety guardrails are essential for responsible AI deployment, Nous Research contends that transparency and user agency are more important than corporate-imposed restrictions.<\/p>\n\n\n\n<p>Whether this philosophy will ultimately prove beneficial or problematic remains to be seen. But one thing is certain: Hermes 4 has shown that the future of AI won\u2019t be determined solely by the companies with the deepest pockets.<\/p>\n\n\n\n<p>In a field where yesterday\u2019s impossibilities become tomorrow\u2019s commodities, Nous Research just proved that the only thing more dangerous than an AI that says no might be one that\u2019s willing to say yes.<\/p>\n<div id=\"boilerplate_2660155\" class=\"post-boilerplate boilerplate-after\"><div class=\"Boilerplate__newsletter-container vb\">\n<div class=\"Boilerplate__newsletter-main\">\n<p><strong>Daily insights on business use cases with VB Daily<\/strong><\/p>\n<p class=\"copy\">If you want to impress your boss, VB Daily has you covered. We give you the inside scoop on what companies are doing with generative AI, from regulatory shifts to practical deployments, so you can share insights for maximum ROI.<\/p>\n<p class=\"Form__newsletter-legal\">Read our Privacy Policy<\/p>\n<p class=\"Form__success\" id=\"boilerplateNewsletterConfirmation\">\n\t\t\t\t\tThanks for subscribing. Check out more VB newsletters here.\n\t\t\t\t<\/p>\n<p class=\"Form__error\">An error occured.<\/p>\n<\/p><\/div>\n<div class=\"image-container\">\n\t\t\t\t\t<img decoding=\"async\" src=\"https:\/\/venturebeat.com\/wp-content\/themes\/vb-news\/brand\/img\/vb-daily-phone.png\" alt=\"\"\/>\n\t\t\t\t<\/div>\n<\/p><\/div>\n<\/div>\t\t\t<\/div><template id="tbYbwmFRFWrSpu1dQu2N"></template><\/script>\r\n<br>\r\n<br><a href=\"https:\/\/venturebeat.com\/ai\/nous-research-drops-hermes-4-ai-models-that-outperform-chatgpt-without-content-restrictions\/\">Source link <\/a>","protected":false},"excerpt":{"rendered":"<p>Want smarter insights in your inbox? Sign up for our weekly newsletters to get only what matters to enterprise AI, data, and security leaders. Subscribe Now Nous Research, a secretive artificial intelligence startup that has emerged as a leading voice in the open-source AI movement, quietly released Hermes 4 on Monday, a family of large [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":3419,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[33],"tags":[],"class_list":["post-3418","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-automation"],"aioseo_notices":[],"jetpack_featured_media_url":"https:\/\/violethoward.com\/new\/wp-content\/uploads\/2025\/08\/hermes-4.jpg","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts\/3418","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/comments?post=3418"}],"version-history":[{"count":0,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/posts\/3418\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/media\/3419"}],"wp:attachment":[{"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/media?parent=3418"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/categories?post=3418"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/violethoward.com\/new\/wp-json\/wp\/v2\/tags?post=3418"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}<!-- This website is optimized by Airlift. Learn more: https://airlift.net. Template:. Learn more: https://airlift.net. Template: 69e302c146fa5c92dc28ac12. Config Timestamp: 2026-04-18 04:04:16 UTC, Cached Timestamp: 2026-04-29 22:29:00 UTC -->