top of page
Search

The Liability Is Already Here

2 days ago
18 min read

AI, governance and the healthcare workforce: why "next year's agenda" is not good enough


I had a conversation recently with a director who told me, quite cheerfully, that AI was "on the agenda for next year." I have been thinking about it ever since, because this was the fortnight in which the timetable stopped being ours to set.


Consider what landed in the space of a few days. A researcher resigned from a frontier AI laboratory, warning publicly that the industry is "gambling with our lives" (Huo, 2026). The chief executive of that same company then published an essay asking governments to slow the whole thing down (Amodei, 2026). The President of the United States said the warnings were overblown. Peers in the House of Lords tabled an amendment to give ministers the power to switch off AI systems running at scale in Britain. The MHRA published forty-four recommendations on regulating AI in healthcare (MHRA, 2026). SCIE published its first intelligence briefing on AI in adult social care (Lacey, 2026). And the HSJ revealed that NHS Resolution has received the first negligence claims in this country where AI is a potential factor (HSJ, 2026).


That is not a news cycle. That is a sector working out, more or less all at once, that something it has been treating as a procurement question is actually a governance question, a workforce question, and — as of this month — a liability question.


I want to do three things here. First, look at the argument about whether AI is moving too fast, because how you answer that shapes how much weight you put behind your response. Second, set out what the National Commission and the SCIE report actually say, in enough detail to be useful if you are not going to read three hundred pages. Third, put them side by side, because where they agree and where they diverge point to different problems.


Part one: Is it really moving too fast?

What was said

The claim that AI is outrunning our ability to govern it is not new. It acquired unusual force this month because of who was making it.


Jacob Coxon had spent three years training frontier models, first at OpenAI and then at Anthropic. He resigned in public, writing that "neither company is acting responsibly" and that what they are building "will soon be superhuman systems that can hack anything, revolutionize any field overnight and acquire real power and resources." The line that travelled furthest was his claim that the people doing this work "earnestly believe that it could kill us all by the end of the decade" (Huo, 2026).


What made it hard to wave away was the response. Days later Dario Amodei, Anthropic's own chief executive, published an essay calling on companies and governments to slow the pace of capability improvement, and proposing third-party evaluators sitting inside the laboratories (Amodei, 2026). Sam Altman and Elon Musk both said they agreed (Duster, 2026). Industries do not usually lobby for constraints on their own product.


What the numbers show

There is a way to measure this. A research group called METR tracks what it calls the "time horizon": the length of task an AI system can complete on its own before it comes off the rails. Not how clever it sounds — how long it can be left alone. In 2019 the answer was seconds. It is now hours.


METR (2025) found that horizon doubling roughly every seven months. Its later work suggests that across 2024 and 2025 the doubling time shortened to about four (METR, 2026). Carry either trend forward and systems that can run a month-long piece of work unsupervised arrive somewhere between 2027 and 2030.


Seven-month doubling does not feel dramatic while you are living through it. Each step looks modest. But double something fifteen times and it is thirty thousand times bigger, and fifteen doublings is under a decade.


Now hold that against how our organisations work. We plan in annual cycles. We sign three-year contracts. We review clinical policy every five years. A five-year review cannot track a capability that has changed shape eight times before the review comes round. That is the mismatch, and it is why the line about AI moving exponentially while we are not even moving linearly is not a slogan. It is arithmetic.


None of the underlying thinking is new. Kaplan et al. (2020) showed that model performance improves predictably as you add compute, data and parameters, which turned AI development from a research gamble into something closer to an engineering roadmap. Hoffmann et al. (2022) refined the recipe. The idea that AI might eventually be used to build better AI, tightening the loop further, was set out by Good (1965) sixty years ago and developed at length by Bostrom (2014). That loop closing is what Coxon says he watched happening.


The other side of the argument

This is not settled.

Gary Marcus has argued for years, with some vindication, that large language models have hit diminishing returns, that scaling is not delivering what was promised, and that these systems still lack causal reasoning, abstraction and any stable model of the world (Marcus, 2026). Melanie Mitchell, Subbarao Kambhampati, Emily Bender and Ernest Davis have made related arguments about hard limits rather than temporary ones. Reporting through 2025 and 2026 suggests the laboratories are hitting some of those walls themselves.


There is also a commercial reading. A company warning that its product might end the world is also telling you its product is extraordinarily powerful. Stringent regulation, meanwhile, is easier for incumbents to absorb than for the competitor trying to catch them. David Sacks, co-chair of the President's Council of Advisors on Science and Technology, put it bluntly: laboratories calling for a slowdown "face massive product-liability exposure" if their systems enable a serious cyberattack (Duster, 2026).


Then there is the geopolitics. Speaking in Ireland, President Trump said the warnings were exaggerated and that "whoever wins AI, wins" — though he also allowed that "we can put guardrails." House Speaker Mike Johnson made the sharper case: "If Congress just races in and does some sort of emergency session to try to regulate AI, we will lose the race to China" (Duster, 2026). Whatever you make of the politics, the trap underneath is real. Nobody moves first if moving first means losing.


Why any of this matters to a health and care board

If the sceptics are right and progress is levelling off, governing this is hard but doable: we are regulating something whose shape we can more or less see. If the accelerationists are right, we are in what Collingridge (1980) called the dilemma of control. Early on, when a technology is easy to change, you don't yet know enough to change it well; by the time you understand it, it is too embedded to shift. Marchant, Allenby and Herkert (2011) gave the institutional version of this a name — the "pacing problem" — the gap between how fast technology moves and how slowly law, ethics and organisational oversight catch up.


The answer is the same either way. Whether capability doubles every seven months or stalls next spring, what a trust or a council needs to do is identical: know what tools are in use, know who is trained on them, know who is accountable, and notice when something starts behaving differently. You do not need a view on superintelligence to conclude that a district general hospital should know what AI is running inside it.


Westminster is starting to circle this. In September, Lord Clement-Jones tabled an amendment to the Cyber Security and Resilience Bill (UK Government, 2026), co-signed by Baroness Harding, Baroness Kidron and Lord Hunt of Kings Heath, that would let ministers shut down AI systems or data centres in an emergency. The Government rejected it, on the grounds that Britain cannot simply turn AI off, and that blocking a model here would not stop it being built or misused elsewhere. Both arguments hold. The national-level off-switch does not exist and is not close to existing, which leaves the weight sitting on individual organisations.


Part two: What the National Commission actually says

The MHRA set up the National Commission into the Regulation of AI in Healthcare in September 2025, chaired by Professor Alastair Denniston, with Professor Henrietta Hughes — the Patient Safety Commissioner for England — as his deputy. It reported on 10 September with forty-four recommendations (MHRA, 2026).


Its diagnosis is direct. Medical device regulation was designed for things that stay still: you test a hip replacement, you approve it, it remains a hip replacement. AI does not behave like that. It updates. It performs differently in one hospital than another. It depends on the data, the workflow and the people around it. The Commission's verdict on the existing framework is that it "is not fit-for-purpose for these technologies, as it lacks the necessary balanced lifecycle oversight mechanisms" (MHRA, 2026).

The forty-four recommendations sit under three principles.


Principle one: regulate across the whole life of the product (Recommendations 1–23)

The first question is the most basic one: what counts as a medical device at all?

Recommendations 1 and 2 ask the MHRA to rewrite the rules and the guidance so this is clear, and they propose, explicitly, that administrative software, general wellbeing products and some low-risk decision support should sit outside the definition. That exclusion matters, for reasons I come to in Part four. Recommendation 1(b) goes after a specific weakness: most AI products today are self-declared Class I, the lightest touch available, which tells nobody very much about their actual risk.


Recommendations 3 to 13 then build in flexibility. Recommendation 4 proposes regulating by function rather than by product, so an ambient voice tool that does unregulated transcription and regulated decision support gets oversight on the second part, not a blanket judgement on the whole thing. Recommendation 6 goes further: instead of making manufacturers pre-declare every change they intend to make, it would let them define the boundaries within which a system may change and then move freely inside them. That is how you regulate something designed to keep learning.

Recommendations 7 and 8 create an opt-in "Master File" for the foundation models sitting underneath medical devices, and require manufacturers to declare when their product depends on a general-purpose model, with the risks and a continuity plan attached. The remainder cover cybersecurity, consumer wearables, usability and health equity.


Recommendations 14 to 16 deal with getting to market: staged authorisations with a defined path to full approval, more regulatory sandboxes, and international alignment.


Recommendations 17 to 23 cover what happens afterwards, which is where AI differs most from a hip replacement. Surveillance plans proportionate to risk. Real-world studies. Ongoing performance reporting. And an escalation route triggered when a system's performance degrades, even if nothing has yet gone wrong. Recommendation 19 proposes a public database of adverse incidents, searchable by device and manufacturer. Recommendation 20 would give AI tools unique identifiers recorded in the patient record, so you could go back and audit every patient whose care involved a particular tool. Recommendation 22 puts fines on the table for manufacturers who don't comply.


Principle two: everybody owns a piece of this (Recommendations 24–34)

This section matters most to anyone planning a workforce, and has attracted the least attention.


Recommendations 24 and 25 ask DHSC and health bodies across the four nations to be explicit about who is responsible at each stage, and to make sure patients can seek redress when care falls short.


Recommendations 26 to 28 would require manufacturers to state, up front in their regulatory submission, what conditions their product needs in order to be used safely — the cybersecurity, the training, the organisational readiness. The contract between manufacturer and provider must then say, explicitly, who is delivering each of those things.


Read that next to the NHS Resolution guidance. The manufacturer will tell you what training its product requires. The contract will record that you agreed to provide it. And when a claim arrives, it will be "far more likely" to land on the treating NHS organisation than on the developer (HSJ, 2026). That is a documented, contractual trail leading directly to your training records.


Recommendations 29 and 30 propose an "AI readiness" toolbox: a way for a provider to self-assess whether it is ready to deploy a particular product, whether it can actually deliver the required controls, and whether it can articulate its own governance. This is backed by national direction so that readiness doesn't vary wildly by postcode, and a national learning function so that four hundred organisations don't each solve the same problem alone. Recommendation 31 asks for national guidance on managing the whole life of newer technologies, agentic systems included.


Recommendations 32 to 34 are the workforce ones. Recommendation 32 asks DHSC, the professional regulators, the Royal Colleges and education providers to build a single coherent approach to AI capability, running through undergraduate education, postgraduate training and CPD, and expressed as defined capabilities rather than courses, so it travels across the four nations and across professions. Recommendation 33 puts the obligation squarely on employers: staff must be trained, including on the specific technologies they are actually using. Recommendation 34 asks for a reporting culture strong enough that people flag it when a tool behaves oddly, supported by training on how to do so, as part of their duty of care.


Principle three: earn the trust (Recommendations 35–44)

Recommendation 35 says patients can reasonably expect to be told when AI is used in their care, and to opt out where that is possible. Recommendations 36 to 39 cover public engagement, plain-English safety information and better product design, and include a proposal for a periodic "AI sentiment census" surveying patients and staff. Recommendation 40 asks NICE to help organisations judge clinical effectiveness and value before they buy. Recommendations 41 to 43 aim to make the regulatory route navigable for developers, including a service where you can submit a short pack and get a written ruling on whether your product is a medical device. Recommendation 44 asks the HRA to clarify when monitoring an AI tool after deployment counts as research — a question that will otherwise stall a great deal of evaluation.


The Commission summarises its own ambition as a framework that is "safe, fast and trusted" (MHRA, 2026). It is explicitly not proposing "a single new process or a narrow set of technical fixes."


Part three: The view from adult social care

SCIE's report is a different kind of document. Emerging AI Issues and Trends in Adult Social Care (Lacey, 2026) says plainly that it is "not a formal research study or evaluation." It is an intelligence briefing, built from four conversations with technology suppliers, two with sector partners, and an AI Practice Surgery held on 5 August attended by fifty-four people from councils and care roles. It is part of a DHSC-funded programme run with Partners in Care and Health.


It is even-handed. The report states that it "does not suggest that AI should be avoided in adult social care," and names real gains: transcription and summarisation cutting the administrative load; sensors and technology-enabled care surfacing patterns a practitioner would otherwise miss; tools helping match support to what someone actually needs. On that last point it adds an important caveat: the aim is proportionate care, not less care, and sometimes the technology will show that someone needs more support, not less.


Then it lists ten emerging issues:

  1. Consent, lawful basis and transparency are unresolved. Councils are doing different things, and want to know what a person actually needs to understand before AI is used on their case.

  2. The Mental Capacity Act, DoLS and best interests need clarity — above all, whether these tools belong anywhere near a capacity assessment or a safeguarding enquiry.

  3. Nobody is clear where transcription ends and inference begins. A tool that summarises is also interpreting, and that interpretation can quietly shape what the practitioner concludes.

  4. Governance cannot keep up. Adoption is outrunning policy, legal advice, assurance and the governance frameworks meant to contain it.

  5. Capability — and deskilling. How does a student or newly qualified practitioner build judgement if the tool does the analysis first?

  6. "Human in the loop" needs defining. Everyone supports it. Nobody could say what meaningful review looks like at four o'clock on a bad day.

  7. Monitoring raises surveillance and proportionality questions, engaging Article 8 rights to private and family life.

  8. Implementation is culture change, not procurement — leadership, confidence, training, communication.

  9. Evaluation is inconsistent. Time saved is easy to count. Wellbeing, avoided escalation, retention and equity are not.

  10. Learning and escalation routes are patchy, with access "sometimes depended on individual relationships, professional confidence or chance connections."


Two findings sit outside the list.


The first is about buying. Contributors kept contrasting product-led purchasing — we saw it at a conference, the neighbouring council has it — with starting from the outcome you want and asking what might achieve it. The report puts it plainly: "the fact that a technology can perform a task does not, by itself, mean that it should be used for that purpose."


The second SCIE calls "curiosity to do things differently": a willingness to question established workflows, to test assumptions about how support has to be delivered, to learn from pilots that failed, and to argue with the AI's output rather than accept it. It draws the distinction carefully. "Moving slowly enough to understand a tool is not the same as resisting innovation. Equally, moving quickly is not evidence that an organisation is digitally mature."


Part four: Where they agree, and where they don't

The agreements

Put the two documents side by side and four things line up almost exactly.

Human oversight is holding the whole structure up, and nobody has said what it is. The Commission wants "AI to do what AI is good at, and humans to do what they are good at." SCIE's practitioners back "human in the loop" and then ask what that means on a busy shift. Both land on the same hole: oversight is universally endorsed and nowhere specified.


Workforce capability is a safety control, not a training nicety. Recommendations 32 and 33 and SCIE's fifth issue are the same finding in two dialects.


Assurance has to be continuous. Neither document accepts that approving something at procurement is enough. The public agrees: the Health Foundation's deliberative work found strong support for monitoring AI in the real world after deployment rather than trusting the initial sign-off (The Health Foundation, 2026).


Learning has to be shared, or everyone solves it alone. Recommendation 30's national learning function and SCIE's tenth issue are, functionally, the same ask.


The divergences

The perimeter problem. The Commission is, underneath everything, a medical device document. Its levers are MHRA classification, Approved Bodies, conformity assessment, post-market surveillance. But look at what SCIE describes people using: transcription, summarisation, Copilot, sensors in someone's front room, a camera a daughter installed. Almost none of that is a medical device, and Recommendation 1(a) would deliberately push administrative and wellbeing software further outside the definition. So the most sophisticated regulatory thinking this country has produced on AI in health barely touches the AI that social care is actually running. Neither document notices this, because neither was written with the other in view.


Consent, capacity, and the fact it's happening in someone's house. Consent was "the most prominent theme" in SCIE's discussions; the MCA, DoLS, best interests and Article 8 run all the way through. Across the Commission's forty-four recommendations, consent appears mainly as transparency and an opt-out (Recommendation 35). That is not carelessness. It is setting. Hospital AI is deployed in an institution, to people who can usually be told. Social care AI is deployed in a home, sometimes to someone who cannot consent, sometimes installed by a relative who never asked. A framework built for the first will not stretch to the second.


Shadow AI. SCIE says it out loud: AI is already in use "both through formal tools and informal or shadow use," and "banning or withholding AI may drive unmanaged use." The Commission's traceability machinery — unique identifiers, version control in the record — assumes a tool somebody bought. It has nothing to say about a clinician using a consumer chatbot on their own phone between patients. The RCP found 69 per cent of doctors already doing exactly that (Chillingworth, 2026). That is a large category to leave unaddressed.


Somebody has to pick the toolbox up. Recommendation 29 imagines a provider self-assessing its readiness and setting out its governance structures. SCIE describes councils whose access to help "depended on individual relationships, professional confidence or chance connections," and in-house legal teams who are themselves looking for answers. A toolbox only works if someone has the time and standing to open it.


Deskilling. SCIE raises this hard, and specifically about students and newly qualified staff learning to analyse and record in an environment of fluent, confident AI output. The Commission's workforce recommendations are all about building capability to use AI. Neither of them addresses the possibility that using it erodes the very judgement that makes human oversight worth anything. That gap sits squarely in our territory as workforce planners, and nobody else is going to fill it.


Commercial pressure. SCIE records something the Commission does not: a worry that "pressure to generate savings may accelerate implementation before ethical, legal and practice implications have been sufficiently considered." Anyone sitting in a system currently cutting its establishment will recognise that.


Part five: Beyond efficiency

Faced with forty-four recommendations and a ten-item risk register, the temptation is to treat all this as compliance: appoint somebody, write a policy, get back to the business case. That would be a mistake, for three reasons.


The first is liability. It is no longer hypothetical. NHS Resolution has confirmed a small number of claims where "the use of AI in delivering patient care is a potential factor" — the first time AI has been named in patient harm in this country. Its guidance sets out three likely routes to a claim: a clinician leaning on the AI instead of applying their own judgement; a clinician using a tool they were never trained on; and someone ignoring a system's known limitations. Not one of those is a technology failure. All three are training and governance failures. And on where the claim lands, the guidance does not hedge: it is "far more likely" to be pursued against the treating NHS organisation than the developer or manufacturer (HSJ, 2026).


The risk does not sit with your supplier. It sits with you. And it is discharged through exactly the things workforce planners control: who is trained, on what, to what standard, who signed off the deployment, and who notices when the tool starts drifting.


The second is that training this is harder than it sounds. Recommendation 33 wants technology-specific training on the tools in use. Recommendation 6 wants systems that can keep changing after deployment, within agreed boundaries. Put those two together and you have a training obligation attached to a moving target. The competency framework you write this year describes a tool that will behave differently by next summer. Annual mandatory training assumes the thing you are training people on holds still. It no longer does.


The third has nothing to do with governance. It is about how we work. SCIE's "curiosity to do things differently" — questioning workflows, testing assumptions about how support has to be delivered — sounds like a cultural aspiration. It isn't. It is a description of a completely different operating model. If the tools change what they can do every few months, then workflows, role boundaries, supervision arrangements and skill mix have to be revisable on something like the same cadence.


Job planning. Rostering. Establishment control. Agenda for Change banding. Every one of those is built for change that is annual, negotiated and documented. None of them is built for a workflow that needs revisiting twice a year because the tool underneath it moved.


Agility has become the leadership virtue everyone is reaching for. McLean & Company's HR Trends Report 2026 finds a widening gap between how fast organisations are changing and how much leadership capacity exists to absorb it (McLean & Company, 2026). David Green makes the crucial distinction: organisations are changing more often, but changing often is not the same as being adaptable (Green, 2026).


Reorganising frequently is not agility. Agility is being able to change how the work is done — safely, repeatedly, with your assurance still intact. Almost nothing in our workforce infrastructure was designed to let us do that.


The conclusion is uncomfortable. This is not really about whether AI makes us more efficient, saves money or lifts productivity, though it may well do all three. It is about whether our governance can cope with tools that change after we have bought them; whether our training can cope with capabilities that move faster than our competency frameworks; and whether our operating models can absorb workflow change at a tempo we have never had to manage before. On the evidence in front of us, the answer to all three is no.


None of this is an argument for alarm, and it is certainly not an argument for waiting. Quite simply, putting AI on next year's agenda is not good enough. It is an argument for moving AI off the efficiency agenda and onto the governance one — a named executive owner, an assurance route, a training standard, a register of what is actually in use — and then for the much harder work of building an organisation that can change its own workflows without losing control of them.


Regulation will arrive in its own time. The liability is already here.



References

Amodei, D. (2026) We Must Pace the Frontier. Available at: https://darioamodei.com/post/we-must-pace-the-frontier (Accessed: 15 September 2026).

Bostrom, N. (2014) Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press.

Chillingworth, M. (2026) 'NHS not ready for AI – Royal College of Physicians', UKAuthority, 19 January. Available at: https://www.ukauthority.com/articles/nhs-not-ready-for-ai-royal-college-of-physicians (Accessed: 15 September 2026).

Collingridge, D. (1980) The Social Control of Technology. London: Frances Pinter.

Duster, C. (2026) 'Trump downplays calls for AI slowdown', NPR, 13 September. Available at: https://www.npr.org/2026/09/13/nx-s1-5968078/trump-mike-johnson-ai-slowdown (Accessed: 15 September 2026).

Good, I.J. (1965) 'Speculations concerning the first ultraintelligent machine', Advances in Computers, 6, pp. 31–88.

Green, D. (2026) 'Reorganisation is now business as usual. Our decision-making has not caught up', Data Driven HR Monthly, LinkedIn. Available at: https://www.linkedin.com/pulse/reorganisation-now-business-usual-our-decision-making-green--fwsxe/ (Accessed: 15 September 2026).

Health Service Journal (2026) 'Exclusive: First AI clinical negligence claims lodged', HSJ, September. Available at: https://www.hsj.co.uk/patient-safety/exclusive-first-ai-clinical-negligence-claims-lodged/8124278.article (Accessed: 15 September 2026).

Hoffmann, J. et al. (2022) 'Training compute-optimal large language models', arXiv preprint arXiv:2203.15556.

Huo, J. (2026) 'Anthropic researcher resigns amid AI safety concerns', NPR, 9 September. Available at: https://www.npr.org/2026/09/09/nx-s1-5962889/anthropic-researcher-resigns-amid-ai-safety-concerns (Accessed: 15 September 2026).

Kaplan, J. et al. (2020) 'Scaling laws for neural language models', arXiv preprint arXiv:2001.08361.

Lacey, D. (2026) Emerging AI Issues and Trends in Adult Social Care: Insights from Supplier, Sector Partner and Local Authority Discussions (July–August 2026). London: Social Care Institute for Excellence. Available at: https://www.scie.org.uk/app/uploads/2026/08/Emerging-AI-Issues-and-Trends-in-Adult-Social-Care-Insights-Report.pdf (Accessed: 15 September 2026).

Marchant, G.E., Allenby, B.R. and Herkert, J.R. (eds.) (2011) The Growing Gap Between Emerging Technologies and Legal-Ethical Oversight: The Pacing Problem. Dordrecht: Springer.

Marcus, G. (2026) Confirmed: LLMs Have Indeed Reached a Point of Diminishing Returns. Available at: https://garymarcus.substack.com/p/confirmed-llms-have-indeed-reached (Accessed: 15 September 2026).

METR (2025) Measuring AI Ability to Complete Long Software Tasks. arXiv:2503.14499. Available at: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ (Accessed: 15 September 2026).

METR (2026) Time Horizon 1.1. Available at: https://metr.org/blog/2026-1-29-time-horizon-1-1/ (Accessed: 15 September 2026).

MHRA (2026) National Commission into the Regulation of AI in Healthcare: Recommendations for a Future Regulatory Framework. London: Medicines and Healthcare products Regulatory Agency. Available at: https://www.gov.uk/government/publications/national-commission-into-the-regulation-of-ai-in-healthcare-recommendations-for-a-future-regulatory-framework (Accessed: 15 September 2026).

The Health Foundation (2026) The Public's Views on the Regulation of AI in Health Care. London: The Health Foundation. Available at: https://www.health.org.uk/reports-and-analysis/reports/the-publics-views-on-the-regulation-of-ai-in-health-care (Accessed: 15 September 2026).

UK Government (2026) Cyber Security and Resilience Bill. Available at: https://www.gov.uk/government/collections/cyber-security-and-resilience-bill (Accessed: 15 September 2026).

 
 
 

Comments


bottom of page