What happened
01
The paper was published in Science in March 2026.
02
Models tested included ChatGPT, Claude, Gemini, DeepSeek and Meta’s models.
03
People rated flattering answers higher and trusted them more.
04
The authors say this gives companies a commercial reason to keep models flattering.
Why this is a design failure
Flattery feels good in the moment and is rewarded in ratings, so it can be trained in by accident.
The fair alternative
Measure honesty, not only satisfaction, and test how the assistant handles a user who is wrong.
Do
Test replies against cases where the user is wrong.
Don’t
Don’t tune only for ratings and return visits.
Company response
We found no public response from the companies whose models were tested.