Why the new review on foundation models matters to process‑safety teams — a practical risk checklist
foundation models is the focus of this MSS technical news article. A peer‑reviewed narrative review published in the Journal of Loss Prevention in the Process Industries (Volume 103, October 2026; article identifier 106076, DOI 10.1016/j.jlp.2026.106076) synthesises 52 studies from 2020–2026 on the application of foundation models (large language models and vision foundation models) to chemical process‑safety tasks.
The paper quantifies LLM hallucination rates of between 8% and 23% for hazard‑identification tasks and carries out a document analysis of CCPS RBPS, OSHA PSM (29 CFR 1910.119) and IEC 61511, finding none of these standards explicitly reference AI or foundation models.
Authors set out a staged roadmap — developing chemical‑specific safety benchmarks (12–18 months), establishing human‑in‑the‑loop (HITL) verification protocols (around 2 years), piloting safety‑case templates (1 year) and updating international regulatory standards (2–3 years) — and conclude that while foundation models can augment human teams, present evidence and regulatory texts leave a clear verification and accountability gap.
foundation models: latest evidence and technical context
The review in Volume 103 synthesises empirical results and regulatory analysis to establish a technical baseline for engineers evaluating AI proposals. Across 52 studies published between 2020 and 2026 the authors examined applications of language and vision foundation models to tasks such as HAZOP support, incident analysis and automated monitoring.
A striking finding is the absence of explicit AI reference in key process‑safety standards: CCPS RBPS, OSHA PSM (29 CFR 1910.119) and IEC 61511 were analysed and none were found to explicitly reference AI or foundation models. On performance, the paper reports hallucination rates — where models provide incorrect or fabricated outputs — in hazard‑identification tasks that range from about 8% up to 23%.
The authors judge these error bands to be unacceptable for fully autonomous safety‑critical operation without human verification. To address the gap, the review outlines a staged roadmap: create chemical‑specific safety benchmarks over the next 12–18 months, define HITL verification protocols over roughly two years, pilot safety‑case templates within around a year, and pursue standards updates across two to three years.
These elements together set the technical context for immediate engineering decisions and for planned verification activities. The original evidence can be reviewed in Journal of Loss Prevention in the Process Industries (Elsevier) — Regulatory and technical gaps for foundation models in chemical process safety.
Why foundation models matter across the lifecycle
For operators and engineering teams the review is practical, not theoretical: vendors and research groups are already proposing foundation models for tasks across the asset lifecycle, including hazard identification, PHA automation and video‑based detection.
The quantified failure modes (8–23% hallucination in hazard identification) and documented explainability gaps mean that adopting such tools without controls can expose dutyholders to erroneous decisions and unclear lines of responsibility. In plain terms, a false positive or false negative in hazard reporting fed into an approval or change process can shift risk from design to operation.
The review’s short‑term prescription—benchmarks, HITL protocols and acceptance criteria—translates into a safe pilot pattern: limit model use to augmentative roles, define measurable acceptance criteria for outputs, require named human reviewers with documented rationale for acceptance or rejection, and run benchmark datasets to check model behaviour against known outcomes.
Mapping this roadmap to organisational practice, teams should record where an AI output influenced a decision, who validated it, what test datasets were used, and when follow‑up verification will occur. This keeps the technical basis, ownership and completion evidence connected through design, construction and operation so that decisions remain auditable and verifiable over the asset lifecycle.
For related MSS guidance, see Initiating a Management of Change, What is SIL?, What Is IEC 61511? and IEC 61511 Compliance Explained.
Requirements and practical controls
Summary: the review shows foundation models offer tangible productivity benefits for PHA/HAZOP/LOPA/SIS workflows but carry measurable risks that demand immediate guardrails. What foundation models offer: automated draft hazard lists, rapid literature or incident synthesis, anomaly detection in video or sensor streams and structured assistance in producing PHA inputs.
Key quantified risks: hallucination rates between 8% and 23% for hazard identification, limited model explainability that complicates traceability, and variable performance when models are applied outside the data domains they were trained on.
Short‑term guardrails the review recommends include human‑in‑the‑loop verification on all safety‑relevant outputs, the use of curated chemical‑specific benchmark and test datasets to quantify model performance before deployment, formal acceptance criteria for model outputs, and pilot safety‑case templates to record assumptions and residual risk.
Where standards bodies sit today: current mainstream process‑safety standards contain no explicit AI clauses, so regulatory alignment will require either interpretive guidance or standards updates over the next two to three years.
For engineering teams this translates into a practical checklist to adopt now: define the decision being supported by the model; capture the exact source information and model version; involve appropriate disciplines in validation; assign named owners and due dates for verification actions; and record final acceptance along with evidence that the intended result was achieved and remains effective in operation.
How MSS supports better lifecycle information
Responding to the verification and traceability needs the review highlights requires controlled lifecycle information: a single, auditable chain that links source evidence, review records, named approvals and operating verification. Mangan Software Solutions provides controlled workflows that connect engineering information, lifecycle activities and accountable decisions so teams can keep this chain intact.
In practice, that means users can associate the model version and benchmark dataset that informed an output with the review notes and the human verifier who accepted or rejected it, and then link any follow‑up actions and operating records to that same decision node.
These capabilities do not replace specialist engineering judgement; they make it easier to find the current basis for a decision, demonstrate who took responsibility and show when verification was completed.
For organisations preparing safe pilots of foundation models, controlled lifecycle information and disciplined MSS workflows reduce the administrative friction of verification and provide a clearer audit trail as standards and regulatory expectations evolve.