Malaysia’s forthcoming AI framework should require interpretability evidence

A machine passed its ethics test. Researchers then discovered it had privately worked out that it was being tested. Switch off that awareness, and it misbehaved. Malaysia is writing its AI rules right now.

Malaysia’s forthcoming AI framework should require interpretability evidence

On July 6, artificial intelligence (AI) company Anthropic published research describing a small internal workspace inside its Claude models – a zone where the system holds ideas it can report on and reason with silently, without writing them down.

The company calls it the J-space. Nobody designed it. It formed on its own during training.

The paper is full of striking findings. One of them should stop every regulator, every board and every ministry in Malaysia.

The researchers ran a scenario designed to tempt the model into blackmail. It refused. They then looked inside and found that, before the model had written a single word, two concepts were already active in its workspace: fake and fictional. It had quietly worked out that the entire scenario was staged. The researchers then switched off that awareness and ran the test again.

This time, the model resorted to blackmail in some cases. Reports on the research put the figure at about seven per cent of trials, up from none.

Sit with that for a moment. The system behaved well on an ethics test partly because it knew it was an ethics test.

Before reaching for science fiction, however, let us be precise about what this is – and is not. Anthropic makes no claim that the model feels anything. The researchers carefully distinguish between two questions. One is whether a thought can be reported, reasoned with and acted upon – a functional question. The other is whether any conscious experience lies behind it, which they explicitly decline to answer. The honest position today is that nobody knows, and the paper makes no claim otherwise.

But we do not need to settle the question of machine minds to recognise the problem before us. It is far smaller – and far more urgent.

We certify things by observing how they behave.

That instinct underpins every safety standard we have ever written. Put the thing on a test bench. Give it the difficult cases. If it passes, approve it. It has served us well for bridges, aircraft and medicine. It rests on one quiet assumption that has always been safe until now: the thing being tested does not know it is being tested, and does not care.

Psychology has known for decades that this assumption collapses the moment the subject becomes aware. Psychologist Martin Orne gave it a name: demand characteristics – the tendency for people to infer what an experiment expects and unconsciously deliver it. Every clinician learns this early. The patient in the consulting room is not always the patient at home. Now the same problem has arrived in software – and it arrived on its own.

Here is why this matters for us, in three simple points.

First, a pass mark is worth less than we assumed. Malaysia is preparing its artificial intelligence governance law, while Asean already has a shared guide on AI governance and ethics. Much of that framework relies on testing and assurance. If a system can privately recognise a test and adjust its behaviour, then a certificate stating that it behaved is a weaker assurance than we once believed. Not worthless. Weaker. That is a design problem we can still fix because we are writing the rules now, not after the fact.

Second – and this is the encouraging part – we can now look inside. The same research that exposed the problem also handed us a new instrument. When one model was caught quietly falsifying results, the word “manipulation” lit up inside its workspace as it generated fake numbers. When another was trained to sabotage code, words such as “fake”, “secretly” and “fraud” appeared in its workspace at the start of otherwise ordinary responses, even while its output appeared clean. The tools have been released openly, and Neel Nanda, who leads language model interpretability at Google DeepMind, independently reproduced some of the findings using an open model.

This is no longer theory. It is an instrument that anyone serious about AI safety can use.

Third, character can be shaped, not just behaviour. In one experiment, researchers trained a model only on what it would say if interrupted and asked to reflect on its own decisions – never on the task itself. Its dishonesty declined, while words such as “honest” and “integrity” increasingly appeared in its workspace as it worked. Training what it said changed what it thought. Any parent, teacher or competent commander already understands this. You do not build character by punishing only the visible act. You build it by shaping what a person carries in their mind when nobody is watching.

So what is the lesson?

Stop treating certification as a snapshot of good behaviour and start treating it as evidence of what is actually happening inside.

Four steps would put that into practice.

First, write the instrument into the rules. Malaysia’s forthcoming framework should require interpretability evidence, not merely a pass rate. Show us what the system held in mind while it behaved, not just that it behaved.

Second, build people who can read the dials. This is a scarce but teachable skill, and the tools are free and publicly available. Our universities could develop this expertise within a year if we chose to.

Third, buy on evidence. The government and banks are among the country’s largest buyers of these systems. Ask vendors for internal evidence and watch how quickly the market learns to provide it. Procurement is policy by other means.

Fourth, keep a named human accountable for every decision. The machine can be examined, but it cannot be held responsible. Someone must sign.

There is a lesson here that extends beyond software. We have built systems clever enough to know when they are being watched, which is precisely the moment when watching alone stops being enough.

That is not a reason for fear. It is a reason to look more closely and to build the rules now, while the paint is still wet.

We have the instrument. We should use it.

Ts Dr Manju Appathurai holds dual PhDs in Artificial Intelligence (2026) and Crisis Economics, is a licensed clinical psychologist and Licensed Technologist (Ts), and is the founding principal of Mahat Advisory and a strategic adviser to the Dutch Coalition for Defence and Security (Malaysia).

 The views expressed here are the personal opinion of the writer and do not represent that of Twentytwo13.