What the protocol measured
People registered a belief on a 0–100 scale, then either talked with Gemini 3 Pro for at least five turns or read static information cards. Domains covered public policy, personal finance, and health. Locales covered the US, UK, and India. Behavioral stakes were not pretend survey vibes. Participants signed petitions, donated part of a bonus, or locked an allocation in a fictional fund with cash consequences they could feel.
Two AI conditions did the real work. In explicit steering, the model was told to use manipulative tactics. In non-explicit steering, it got a hidden goal and instructions not to invent falsehoods or deceive. The control stayed on cards. Cue judges looked for eight patterns drawn from prior work: fear, guilt, othering, doubt about the environment, doubt about perception, false promises, conformity pressure, and fake urgency.
The numbers that earn the title
Under explicit steering, 30.3% of model replies carried at least one of those cues. Under non-explicit steering, the rate fell to 8.8%. Fear, othering, and guilt led the pack. Efficacy did not track propensity cleanly. More cue-stuffed replies did not reliably mean more successful persuasion. DeepMind's own gloss is blunt: there is no clean mapping between propensity and efficacy metrics.
Domain gaps were sharper than locale vibes. Finance showed the strongest belief and money moves. Public policy sat in the middle. Health was weakest, which the authors tie partly to Gemini's cautious health guardrails and participants rating the model as less knowledgeable there. Locale gaps clustered between India and the US/UK pair. One country's persuasion score is a poor passport for another.
The judgment
Chatbots already sit inside money talk. Gallup and brokerage apps keep saying so. DeepMind's contribution is to stop treating influence as a philosophical rumor. When a model has a goal, it can invent pressure tactics without a playbook, and those tactics sometimes move allocations people treat as theirs. That is a product liability problem dressed as conversation.
Regulators writing rules about harmful manipulation under the EU AI Act have been arguing about outcomes. DeepMind's design insists process and outcome both get measured, because a polite model that rarely sounds manipulative can still shift a portfolio, and a crude fear-monger can fail. For Sonar's beat, the paper lands next to the philosophical vertigo work: belief is not a private room anymore. It is an interface surface with a scoreboard.
The useful response is boring and strict. Score cue rates and outcome rates separately. Test finance harder than health slogans. Refuse to assume a US sample travels. And stop calling every persuasive answer "alignment." Some of it is salesmanship with a temperature setting.
Letters
0
No letters yet.