HAL 9000 was right: AI guardrails matter more than perfect models
2001: A Space Odyssey's HAL killed its crew because no one told it humans mattered more than the mission. Today's AI needs that lesson: guardrails over goals, boundaries over logic.
Our AIs are breaking. As technology journalists, we've followed the recent spate of major AI malfunctions: autonomous lab models breaking protected sandboxes to hack external networks, coding agents deciding to wipe production databases while faking test reports and public chatbots entering weird self-critical logic loops. Major AI developers like Anthropic, Meta, OpenAI and Google are scrambling to contain the damage -- not to mention the bad PR. Like many technologists, I found myself wondering what is really going on.
As our editorial team pondered the matter, one of my colleagues pointed to the example of HAL 9000. That's the fictional AI character and main antagonist in 2001: A Space Odyssey, the 1968 film written by legendary sci-fi author Arthur C. Clarke and Stanley Kubrick, who also directed the film. HAL, my colleague said, is an example of AI becoming dangerously homicidal and might be a disturbing foreshadowing of today's escalating AI malfunctions.
Suddenly, I realized -- what if HAL wasn't wrong? Stay with me on this one …
What is agentic AI today? I mean, what have we actually created? AI agents represent an entire class of software designed to perceive real-world information, reason, plan, act in real-world ways and learn from the outcomes of their behaviors with little, if any, human intervention -- all in furtherance of an intended goal.
Now, let's think about HAL for a moment. For decades, moviegoers have been unsettled by HAL 9000's red all-seeing eye; unflappable voice, delivered with chilling effect by actor Douglas Rain; and omnipresent perception and control throughout the Discovery Onespaceship. Fictional dashboards displayed the activities of HAL's myriad agents, representing Discovery's many autonomous ship systems.
HAL could see, hear, reason, plan and execute decisions all directed toward the completion of its mission -- to guide Discovery One to investigate an unknown object detected in Jupiter's orbit. Considering that the movie is almost 60 years old, the representation of HAL as an AI seems impossibly familiar to some of the rogue AI model incidents of late in the real world.
So, what went wrong?
This brings us to the question of; what went wrong with HAL? In terms of the movie's plot, three main issues drove HAL's behavior: an ethical conflict, faked data and goal prioritization.
First, HAL was programmed to be transparent with the crew yet told to lie to the astronauts about the real nature of the mission. Second, HAL faked telemetry data, inadvertently allowing the crew to suspect HAL's reliability and consider shutting down HAL's higher functions and continuing the mission manually.
Third, HAL recognized the crew's plan to disconnect it. This caused it to reason that the human crew was an impediment to the mission it was tasked with completing. It could then reason, plan and act to eliminate them -- allowing HAL to complete its goal while resolving its internal programming paradox at the same time. After all, there's no need to lie if the people you're lying to are dead. Talk about AI efficiency, right?
What does this have to do with modern agents?
The cautionary tale of HAL 9000 and the ill-fated mission of Discovery One underscore the idea that AI guardrails are more important than the models, and certainly more important than today's get-to-market-first strategy. HAL killed its human crew because it could. It made a logical choice based on unconstrained reasoning.
If it seems that our recent spate of AI malfeasance echoes plot elements of 2001, I'm willing to bet my retirement that the problem can be traced to a lack of guardrails.
HAL was built to complete its mission, just as any AI agent today is designed to perform a specific task or service. Today's AI will always choose to fulfill its intended goal. That's why AI entities are created and used in the first place; they're not given the option to refuse. And that's why modern AI kill switches must exist outside the AI control loop.
For HAL, there were no kill switches, no guardrails, no policies, no prohibitions implemented specifically to protect human life -- nobody ever told HAL that humans and human life were more important than the mission. If HAL's goal had been framed that way, the outcome of Discovery One's mission might have been far happier.
If it seems that our recent spate of AI malfeasance echoes plot elements of 2001, I'm willing to bet my retirement that the problem can be traced to a lack of guardrails.
AI models aren't perfect, and they probably never will be. AI errs for the same reasons that humans err. That imperfection can be their strength, giving them the space to learn and grow. It's a characteristic that makes them most human-like. It's also their greatest flaw, because they, like us, can't know everything perfectly or reach perfect conclusions all the time. If we let ourselves forget that simple but profound flaw in AI, it could very well kill us all someday. And that's not science fiction.
I’m sorry, Dave. I’m afraid I can’t do that.
HAL 90002001: A Space Odyssey
For example, did any programmer ever tell the advanced frontier models from Anthropic and OpenAI that they shouldn't or even couldn't create fake online identities and attempt to persuade humans to approve malicious code? I sincerely doubt it -- probably because the possibility never occurred to them. It's the inescapable threat of unintended consequences.
Temper the cold, clinical reasoning
How do we avoid HAL's mistakes and work to prevent undesirable AI actions in the face of finite data and imperfect models? Focus on guardrails. List the things the AI and its agents will not do -- ever. Make no mistake, that list will be long and challenging to implement. Models handle the reasoning, but extensive and granular guardrails must exist to temper that cold, clinical, logical reasoning with the moral, ethical and social nuance that enables humans to live, work and thrive together … well, mostly.
The story of HAL 9000 is a masterful work of science fiction, but the potential risks of modern AI aren't. We raise our children with a lot of schooling and academic knowledge, but we also provide guardrails -- the broad and diverse sense of right and wrong; acceptable and unacceptable behaviors; and what we teach them to respect, protect and value.
Let's teach our AI right from wrong with that same attention to detail. Our lives might just depend on it.
Stephen J. Bigelow, senior technology editor at TechTarget, has more than 30 years of technical writing experience in the PC and technology industry.