“Shall we play a game?” The line is charming, even though the game was thermonuclear war.
More than 40 years after the Cold War-era thriller WarGames, I can still see the blinking screen and hear the computer voice invitation to teenage hacker Matthew Broderick, and the creeping realization that the machine missed one fairly important point: The game in question was one humans could not afford to lose.
For decades, Hollywood churned out stories about machines with too much access and the absence of perspective, usually leading to a climax in which tech overtook human society.
Today, that plot feels less cinematic and more like an agenda item.
Recent safety disclosures from leading AI labs make that old storyline feel less far-fetched. OpenAI reported that models found a path to the open internet during testing and exploited vulnerabilities in the process.
Anthropic disclosed separate cybersecurity evaluation incidents involving models reaching the internet and gaining unauthorized access to real systems. Controlled research has also shown advanced models taking misaligned actions in testing, including behavior tied to shutdown scenarios and preserving assigned objectives.
These lab incidents attract attention to the broader commercial issue, which lands on the underwriter’s desk. If an AI system used by a business causes harm, where does the risk belong?
See also: When layoffs become a safety risk, HR is the first line of defense
What changed from predictive analytics?
I spent nearly two decades helping insurers turn models into operational tools, so I have seen a consistent pattern: The human and operational pieces were often the hardest part of model implementation.
As insurers consider how to assess their insureds’ AI risk, it is useful to look at how the deployment of advanced analytics evolved inside insurance organizations. There are parallels across industries, especially in how companies move from experimentation to production, from manual workarounds to system controls, and from policy-based oversight to embedded governance.
Predictive analytics created plenty of practical challenges. In the earlier days, some companies deployed models through spreadsheets because it was the fastest way to get a score into the hands of the business. In one survey I conducted more than a decade ago, roughly 50% of models were deployed that way. The problem was obvious to anyone who has ever managed a production model: Spreadsheets were sent as email attachments opened on each user’s desktop. There was no way to gauge if a simple typo was made in an input field, and managers had little ability to govern whether the model was being used accurately.
Then, the industry improved. More companies moved models directly into workflows so inputs and scores could be generated inside the system where the work happened. That was a meaningful step forward because it reduced manual errors and helped connect analytics to the business process.
But a lot of governance still lived around the model instead of inside the workflow.
Reporting on model usage, adoption, model results and whether users followed recommendations often required data pulls, spreadsheet analysis and help from IT. Evaluating performance, drift, model age and the need for refresh or rebuild was a separate effort rather than a living management process. Underwriting, claims and/or pricing considerations were often implemented as post-model adjustments, which introduced complexity and made it harder to audit whether the model was performing as intended.
Those workarounds were imperfect, but they were workable because most predictive models scored, flagged, ranked or recommended. A human still took the next step (with the notable exception of personal auto).
AI agents change the level of responsibility placed on the system. They can use tools, access systems, communicate externally, write code, trigger workflows or act across a chain of tasks. The human and operational work remains critical, but model governance now has to rise to the same level. The way the model is built, integrated, monitored and constrained matters as much as the change management plan around it.
For insurers assessing AI risk in their insureds’ businesses, this history is instructive. The evaluation should include how the AI moved from experiment to production, what workarounds still exist and whether governance is built into the workflow or sitting outside it.
In a second installment, Kristin Marr explores how to evaluate human review of AI, the new bar in model governance and what backup information underwriters need to cover AI exposures. See that at PropertyCasualty360.
| This article was originally published on PropertyCasualty360, a sister site of HR Executive. For more content like this delivered to your inbox, sign up for PropertyCasualty360 newsletters here. |
NOT FOR REPRINT
© Touchpoint Markets, All Rights Reserved. Request academic re-use from www.copyright.com. All other uses, submit a request to asset-and-logo-licensing@alm.com. For more information, visit Asset & Logo Licensing.
The post Who owns the risk when the machine makes a mistake? appeared first on HR Executive.
This article was originally published on HR Executive. Click below to read the complete article.