Exhibitors & Products
Events & Speakers
Daily Program

The two go hand in hand – with the result that AI laboratories are poaching humanities graduates from taxi companies, because the characteristics of their models can no longer be determined by technical factors alone. According to a February analysis by the Federal Reserve Bank of New York, the unemployment rate among philosophy graduates is now lower than that of computer science graduates.

In May, Google DeepMind recruited the Cambridge philosopher Henry Shevlin, previously deputy director of the Leverhulme Centre for the Future of Intelligence. The company described the role as its first official in-house philosopher position. His remit covers machine consciousness, the relationship between humans and systems, and what DeepMind refers to as ‘AGI readiness’. Shortly afterwards, Atoosa Kasirzadeh – who holds a PhD in the philosophy of science and mathematics and was previously at Carnegie Mellon University – joined the company.

At Anthropic, at least two philosophers – Amanda Askell and Joe Carlsmith – are part of the team that drafted the constitution for the Claude model. Sam Altman has stated that OpenAI consulted hundreds of moral philosophers when defining ChatGPT’s rules of conduct – yet evidently none of them was able to convey to Altman the ethical value of open models.

What the labs are actually buying

The obvious attempt at an explanation is: the companies want to ease their conscience. The more precise one is: they have a specification problem.

A language model makes decisions every second that cannot be derived from training data. What does a system do if an instruction from the operator contradicts an instruction from the user? When does it refuse to provide information, and where does the line between caution and uselessness lie? And could the AI system have an intrinsic motivation? What would that look like? These are not questions that an engineer can answer with a test case. They are, first and foremost, questions of priority rules between competing demands and, to a large extent, questions of meaning; and for this spectrum, there has been a discipline in existence for at least two and a half thousand years.

Several laboratories are now investigating whether and when their systems are entitled to moral status. Anthropic created a dedicated role for this in 2024 and built a team around it. The answer to this question remains open, and it does not fall within the remit of technology.

Putting it to the test

Whether such roles are tenable became apparent a few days ago. Jacob Coxon, a 27-year-old British man and trained mathematician responsible at Anthropic for training new models using very large datasets, not only announced his resignation publicly that day, but also his departure from the entire industry. He had only moved from OpenAI to Anthropic this year, precisely because the company is known for its safety research.

Speaking to the Wall Street Journal, which first reported the story, Coxon explained that he was taking this step because he no longer wished to be part of a race to develop systems that improve themselves. He told the newspaper that the industry was on a course where “by the end of next year, things could already be out of control”. Safety compromises were inevitable as long as the companies were competing against one another; in his circle, terms such as ‘crunchtime’ and ‘endgame’ were now being bandied about. In a post on social media, he wrote, ‘no other human activity carries this level of risk’, and neither his former nor his current employer was acting responsibly.

According to the Wall Street Journal, Coxon is the first Anthropic employee to leave for safety reasons; there have been several similar cases at OpenAI in recent years. On the same day, the head of alignment research at Anthropic publicly estimated the risk of humanity’s extinction by AI in the coming decade at over ten per cent.

Dealing responsibly with something inexplicably powerful

This also highlights the objection to the new roles, and it comes from the right source: Edward Harcourt, director of the Institute for Ethics in AI at Oxford, has pointed out the risk of ‘ethics washing’. It benefits a technology company’s public image if the impression is created that something inexplicably powerful is being handled responsibly. A handful of philosophers can enhance a brand’s image without changing a single product decision. A company that employs philosophers and whose experts are publicly jumping ship either has a functioning internal debate or a structural problem. From the outside, it is difficult to tell the two apart.

Why this is also a procurement issue

For industrial companies that operate on a more economic and less philosophical basis, the crucial point is perhaps not who does the thinking, but where the result ends up. The normative specifications drawn up by these teams are not mere footnotes; they are product features. They are set out in model definitions, behavioural specifications and system cards; they determine which requests an assistant processes and which it rejects, which instruction takes precedence, and how an agent makes decisions when faced with conflicting objectives.

To date, these documents have rarely been read by the procurement department. This should change, for one simple reason: in an agent-based system, the priority rule between competing instructions is not an ethical afterthought, but the operational logic. Anyone unfamiliar with it cannot predict the system’s behaviour in the event of a conflict and, consequently, cannot ensure its reliability. And a supplier whose own specialists are leaving the industry is, in this context, just another metric among supplier performance indicators.

Questions whose answers cannot be purchased

The second consequence is the more uncomfortable one. The questions that laboratories are asking themselves arise, on a smaller scale, in every organisation that deploys agents. What decisions is a system permitted to make without seeking further guidance? What happens if an economic requirement conflicts with a safety requirement? Who is liable for a decision that was formally correct but ultimately wrong? And – now it gets almost eerie – what character traits might the AI agent possess? These questions cannot be outsourced to a supplier, as the answers depend on the organisation’s own risk profile, not that of the provider.

Research institutes have hired philosophers for this purpose. Industrial companies may not – yet! – need any, but they do need someone who understands what is at stake. Before the systems go live. In most organisations, this responsibility does not yet exist.

v-cloak>