Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

AI, Safety and What Goes Wrong

The failure modes of AI, stated plainly, with the numbers.

Lesson 57 of 738 min

The systems that decide what you see

The oldest deployed AI, and the largest

Recommendation systems predate the current wave and reach more people than any chatbot. They decide the order of a feed, the next video, the products shown, the news surfaced. They are trained on one broad objective: predict what you will engage with.

That objective is the entire story, and it has a well-understood defect. Engagement is a measurement of your impulses, not of your preferences. Behavioural research distinguishes what people want in the moment from what they endorse on reflection, and the two systematically diverge. A recommender trained on clicks optimises the first and is indifferent to the second, not out of malice but because clicks are what it can measure.

Several consequences follow mechanically.

Emotionally arousing content wins. Anger, outrage and threat produce reliable engagement. No designer chose to promote outrage; outrage won the optimisation.

Extremity is a gradient. Slightly more extreme content is slightly more engaging, and a system taking small steps up a gradient arrives somewhere it was never pointed at.

Uncertainty is unrewarded. Careful, hedged material engages worse than confident material, so the information environment shifts towards confidence regardless of merit.

Feedback closes. What you engage with shapes what you are shown, which shapes what you engage with. This is the loop from the second module, running on your attention.

What the evidence says, honestly

The research is more mixed than either the alarmed or the dismissive account allows, and it is worth stating carefully.

Deactivation experiments — paying people to stop using a platform for weeks — have found improvements in reported wellbeing and substantial time freed. Large collaborative studies with platform data during the 2020 US election found that altering feed algorithms changed what people saw and how much they used the platform, but did not much move political attitudes over the study period. Studies of the "rabbit hole" hypothesis find that recommendation does contribute to exposure to extreme content, while much consumption is also driven by subscriptions and off-platform links.

The defensible summary: these systems clearly shape attention and time, the evidence that they directly change political beliefs at population scale is weaker than commonly assumed, and effects concentrate in some users rather than spreading evenly. Anyone who tells you the science is settled in either direction has not read it.

Generative systems inherit this

The reason this sits in an AI safety course is that the same optimisation is arriving in conversational products, with a difference in kind.

A feed selects from things other people made. A generative system makes the thing, tailored to you, in the moment. If it is optimised on engagement or satisfaction signals, it is a recommender with an unlimited catalogue — able to produce exactly the content that holds this particular person, rather than finding the closest match in a library. That is the sycophancy of the last module and the attachment mechanics of this one, given a business model.

Watch for the signals: streaks, notifications, personas that express missing you, memory used to deepen attachment rather than usefulness, and metrics reported in time spent rather than tasks completed.

What actually works

Self-control advice mostly fails, because it asks you to out-compete a system optimising against your impulses full-time. What works is changing the environment so less willpower is required.

Remove the source of recommendation. Chronological or subscription-only views where available, browser extensions that hide feeds, and following by RSS all replace ranked infinite streams with finite ones. A finite list ends; a ranked stream does not.

Separate discovery from consumption. Collect things to read into one place, and read from there later. This breaks the loop, because what you saved yesterday is a decision your reflective self made.

Set the friction where the impulse is. Log out, remove from the home screen, keep it on one device. Small frictions are effective against impulse and irrelevant to intention, which is exactly the discrimination you want.

Track outputs, not inputs. Measuring hours reduced is a weak intervention. Noticing whether you finished the book, made the thing, saw the people, is a strong one.

And apply the same test to any AI product you adopt: is it optimising for you finishing your task, or for you coming back? The answer is usually visible in the interface within a week.

The one thing to keep

Recommenders optimise engagement, which measures impulse rather than reflective preference, and a generative system with the same objective is a recommender with an unlimited catalogue — so change the environment rather than trying to out-compete it with willpower.

Before you move on

Why is a generative system optimised on engagement structurally more powerful than a feed ranking algorithm optimised the same way?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly