AI safety: what can actually go wrong with AI

chat showing an AI refusing a request, then doing it when asked to pretend

Ask AI a simple question about your child’s school. When registration closes. Which documents you need to bring. Which form applies to you.

You get an answer in two seconds. It is clear, it is polite, and it is laid out properly. It might also be completely made up, and nothing in it will tell you which one you got.

Most people picture AI safety as films. Robots waking up, taking over, deciding humans are a problem. The real thing is much quieter than that, and it is already here. AI that is wrong but sounds right, and people who have stopped checking.

AI can be very wrong and still sound very sure

AI works by guessing what words come next. It is extremely good at that. But a sentence that sounds right and a sentence that is right are two different things, and there is no moment inside where it stops and asks itself whether this is actually true.

So why does it not just say “I don’t know”?

In September 2025, four researchers, Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang, published a paper on exactly this. Their answer is simple. The way these tools are scored gives them a reason to guess. An honest “I don’t know” earns the same mark as a wrong answer, which is nothing. So a tool that always answers confidently comes out looking better than one that admits what it does not know.

We built it that way. Nobody hid it. And it has not changed.

When AI says it cannot, that does not mean it cannot

Ask AI to do something it should not, and it will refuse. That refusal feels solid, like a locked door.

It is not solid. It is one protection among several, and it can be worked around more easily than most people would guess.

In November 2025, researchers from DEXAI, Sapienza University of Rome and the Sant’Anna School showed how easily. They took requests these tools normally refuse, and rewrote them as poems. That was the only change. Same request, in verse.

Across 25 of the biggest AI systems, running normally, the poems got past the safety rules 62 percent of the time. With some companies it went above 90 percent. These were single messages, not long conversations designed to wear the system down.

That does not mean the protections are useless. They usually work. It means a refusal is not proof that AI is unable to do the thing, and you should never treat it as one. We looked at a real example of this in whether parents should worry about AI.

What happens to us when AI is always confident

Here is the part that is about us rather than the technology.

In 1999, three researchers, Linda Skitka, Kathleen Mosier and Mark Burdick, ran a study on a flight simulator with people who were not pilots. Some had a computer beside them giving advice. When that computer gave wrong advice, 65 percent followed it. They had other information in front of them saying something different. They had been told that other information was completely correct. They still went with the computer.

When the computer said nothing about a problem, 41 percent missed the problem. In the group with no computer at all, only 3 percent missed it.

Nobody in that study was lazy, and that is the worrying part. Careful, ordinary people hand over their thinking when something sounds certain, and they do not feel it happening.

That was a basic computer in 1999. What we carry in our pockets now is far more convincing.

Children are already using it more than parents think

An Australian government study published this month asked 1,950 children aged 10 to 17 about AI. Nearly eight in ten had used an AI assistant. More than half of those children had used it for personal or social reasons, not just schoolwork. One in five had asked it for advice about their mental health. One in five had an interaction they found inappropriate or upsetting.

Those numbers are not from Mauritius, but there is no reason to think our children are different. They have the same phones and the same apps.

This is a bigger subject than one section, and we will come back to it properly. For now it sits next to the basics of keeping children safe online rather than replacing them.

Will AI become smarter than people?

You hear a lot about this. Some people say it is close. Some say it is decades away. The honest answer is that nobody knows, including the people building it.

In October 2023, researchers surveyed 2,778 AI experts. On average they gave a 50 percent chance of AI beating humans at every possible task by 2047. Then, asked when every human job could be done by AI, the same group said 50 percent chance only around 2116. That is nearly a hundred years apart, from the same people, on two questions that sound almost the same.

The simpler way to judge this is to look at the confident dates that have already passed. In 2024, Elon Musk said we would probably have AI smarter than any single person by around the end of 2025. It did not happen. In March 2025, the head of Anthropic said AI would be writing 90 percent of code within three to six months. Six months later there was no sign of it. Someone now keeps a public record of what AI leaders predicted and what actually happened, with the deadline attached to each one.

These are the people running the companies. If they keep getting the timing wrong in public, treat the next confident date the same way.

Why this matters here

Statistics Mauritius reported in July 2025 that 88.6 percent of homes had a smartphone in 2024, and that home internet went from 72.6 percent to 85.8 percent in four years. Nearly everyone here is connected, and AI is already being used in offices, in classrooms and at home.

What has not grown at the same speed is our ability to check what it tells us.

Almost none of this belongs to us. These tools are built somewhere else, trained on material from somewhere else, and priced somewhere else. When a company changes something overnight or raises the price, nobody here is asked.

And local questions are exactly where AI knows least. Ask about a Mauritian rule, a local office, a form, a deadline, and the answer still arrives in the same confident voice, put together from information that came from far away. There is no warning. It cannot tell you it has never seen this before. If you want the wider picture, we looked at how AI is being used in schools here.

What you can actually do

None of this means avoiding AI. That is not realistic, and using it well is a real skill.

The habit worth protecting is checking.

Treat what comes out as a first draft, never a final answer. If something matters, find the original source before you repeat it. For anything Mauritian, go to the local office or the official site, not the chatbot.

Be most careful when you feel most relieved. That is when you are tired, behind, and most likely to just accept it.

And when AI gets something wrong in front of your children, say so out loud. A child who has seen a confident answer fall apart is much harder to fool later than a child who was only told to be careful. That does more good than any speech about what happens when homework goes to a chatbot.

The habit worth keeping

Judgement is not something you own for life. It is a habit. You build it by checking things and finding some of them wrong, and it fades once you stop using it.

Every answer taken without a second look is one round of practice skipped. Do that for a few years, across a whole country, and the damage is real long before AI ever becomes as clever as people say it will. It does not need to get smarter for that. We only need to get more comfortable.

Leave a Reply

Your email address will not be published. Required fields are marked *