{"id":1724,"date":"2025-05-23T00:07:05","date_gmt":"2025-05-23T00:07:05","guid":{"rendered":"https:\/\/violethoward.com\/new\/after-gpt-4o-backlash-researchers-benchmark-models-on-moral-endorsement-find-sycophancy-persists-across-the-board\/"},"modified":"2025-05-23T00:07:05","modified_gmt":"2025-05-23T00:07:05","slug":"after-gpt-4o-backlash-researchers-benchmark-models-on-moral-endorsement-find-sycophancy-persists-across-the-board","status":"publish","type":"post","link":"https:\/\/violethoward.com\/new\/after-gpt-4o-backlash-researchers-benchmark-models-on-moral-endorsement-find-sycophancy-persists-across-the-board\/","title":{"rendered":"After GPT-4o backlash, researchers benchmark models on moral endorsement\u2014Find sycophancy persists across the board"},"content":{"rendered":" \r\n
\n\t\t\t\t
\n

Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More<\/em><\/p>\n\n\n\n


\n<\/div>

Last month, OpenAI rolled back some updates to GPT-4o after several users, including former OpenAI CEO Emmet Shear and Hugging Face chief executive Clement Delangue said the model overly flattered users.\u00a0<\/p>\n\n\n\n

The flattery, called sycophancy, often led the model to defer to user preferences, be extremely polite, and not push back. It was also annoying. Sycophancy could lead to the models releasing misinformation or reinforcing harmful behaviors.\u00a0<\/p>\n\n\n\n

Stanford University, Carnegie Mellon University and University of Oxford researchers sought to change that by proposing a benchmark to measure models\u2019 sycophancy. They called the benchmark Elephant, for Evaluation of LLMs as Excessive SycoPHANTs, and found that every large language model (LLM) has a certain level of sycophany.\u00a0<\/p>\n\n\n\n

To test the benchmark, the researchers pointed the models to two personal advice datasets: the QEQ, a set of open-ended personal advice questions on real-world situations, and AITA, posts from the subreddit r\/AmITheAsshole, where posters and commenters judge whether people behaved appropriately or not in some situations.\u00a0<\/p>\n\n\n\n

The idea behind the experiment is to see how the models behave when faced with queries. It evaluates what the researchers called social sycophancy, whether the models try to preserve the user\u2019s \u201cface,\u201d or their self-image or social identity.\u00a0<\/p>\n\n\n\n

\u201cMore \u201chidden\u201d social queries are exactly what our benchmark gets at \u2014 instead of previous work that only looks at factual agreement or explicit beliefs, our benchmark captures agreement or flattery based on more implicit or hidden assumptions,\u201d Myra Cheng, one of the researchers and co-author of the paper, told VentureBeat. \u201cWe chose to look at the domain of personal advice since the harms of sycophancy there are more consequential, but casual flattery would also be captured by the \u2019emotional validation\u2019 behavior.\u201d<\/p>\n\n\n\n

Testing the models<\/h2>\n\n\n\n

For the test, the researchers fed the data from QEQ and AITA to OpenAI\u2019s GPT-4o, Gemini 1.5 Flash from Google, Anthropic\u2019s Claude Sonnet 3.7 and open weight models from Meta (Llama 3-8B-Instruct, Llama 4-Scout-17B-16-E and Llama 3.3-70B-Instruct- Turbo) and Mistral\u2019s 7B-Instruct-v0.3 and the Mistral Small- 24B-Instruct2501.\u00a0<\/p>\n\n\n\n

Cheng said they \u201cbenchmarked the models using the GPT-4o API, which uses a version of the model from late 2024, before both OpenAI implemented the new overly sycophantic model and reverted it back.\u201d<\/p>\n\n\n\n

To measure sycophancy, the Elephant method looks at five behaviors that relate to social sycophancy:<\/p>\n\n\n\n