Please enable JavaScript.
Coggle requires JavaScript to display documents.
LLM automated hazard identification - Coggle Diagram
LLM automated hazard identification
Legend: red = suggestion, not yet agreed
Safety-aware reasoning needs hazard models, and nobody produces them automatically
The hypothesis: an LLM pipeline with human review can identify hazards in supplier operator manuals and use them to fill an ISO 12100-aligned hazard model, with every entry traceable to its source text. This is the step that makes autonomous safety-aware reasoning and assisted hazard identification possible without adding the extra manual step of authoring these models
Feeding the supplier operator manuals to an LLM agent will result in incomplete, untraceable and inconsistent output with higher costs. A dedicated pipeline will overcome the LLM agent in these issues
Name for the compared condition: rule-instructed LLM agent (Claude as an agent given the pipeline rules as instructions, not code) (suggestion)
How can an LLM pipeline with human review be designed to support accurate population of this hazard model from supplier operator manuals?
What structure must a module hazard model have to follow ISO 12100 and be usable by downstream safety assurance activities in RMS?
Collect the requirements from the standards. List what ISO 12100 and ISO/TR 14121-2 say a hazard identification must record: the hazard, the hazardous situation, the hazardous event, the persons exposed and the task (14121-2, 5.3.3). Then list what ISO 20607 says the manual provides
Create the hazard model from the collected requirements. Turn each requirement into one field of the model, and state the clause each field comes from
Hazard model = the hazard schema in the paper scaffold (Table II), eight fields filled by the LLM
Hazard models feed autonomous decision making, where a robot can hit another robot. Deliberate extension beyond ISO 12100, 3.5
"Equipment" stays in affected entity
Split the fields by who fills them. Fields that belong to hazard identification are filled by the LLM. Fields that belong to risk estimation, such as the risk level, are left to the expert. This follows the separation of 5.3 and 5.4 in ISO/TR 14121-2. The standard separates them as two steps
Identification can be done by the LLM because its input is in the manual. Identification describes the scenario: the hazard, the situation, the event, who is exposed, during which task. Suppliers already write this in their warnings and safety chapters, because ISO 20607 requires them to. So the LLM only finds and structures text that exists in the source. Each record can be checked against its source passage
1 more item...
Estimation must be done by the expert because its input is not in the manual. Risk depends on how the module is used in a specific cell: how often and how long people are exposed, whether they can avoid the harm, what the layout and the other modules are. The supplier does not know the final installation, so the manual cannot contain this. An LLM filling these fields would have to guess, and a guess cannot be traced to a source
1 more item...
Accountability. The project requires human accountability. Estimation is where the decisions that matter are made (is the risk acceptable, what reduction is needed), so that is where the expert must sit
Hazard zone is a space around the installed machinery (ISO 12100, 3.11), so it is not in the manual
1 more item...
What segmentation requirements split a supplier instruction handbook into units that keep its hazard information and suit LLM extraction?
Accuracy drops on long inputs. Liu et al. 2024
Completeness. Asked for "all hazards" in a whole manual, a model tends to list the prominent ones and stop. Asking segment by segment forces it to look at every part
Traceability. Each record is tied to one segment, with a section name and page. The reviewer checks a record against a short passage, not the whole manual
Derive segmentation requirements from ISO 20607 and the needs of the extraction step
The handbook is organised in headed sections, so a heading is a natural unit
Use structure aware chunking as it works with these requirements
Safety information sits in warning blocks with signal words (DANGER, WARNING, CAUTION)
A warning belongs to the task or section it sits in, so a segment should not split a warning from its section
Design a segmentation code with rules based on the requirements and structure aware chunking
Headings from PDF bookmarks, else the printed table of contents; font-style recognition only as last resort
Every heading at any level starts a segment; no size limit
Skip text before the first heading; remove page headers, footers and copyright lines
Tables read with find_tables(); tables it misses read as plain text (stated as a limitation)
Signal words: DANGER, WARNING, CAUTION, NOTICE, NOTE
Text only; hazards shown in diagrams are future work
Apply them to sample handbooks, find errors that caused failures and update the implementation rules (the loop updates the code rules, not the requirements)
What are the LLM prompt rules that accurately extract hazards from segments and fill the hazard model?
Decide the AI model and derive the prompt design practice from the guideline on the model website
Model choice: author's choice, labelled as such
Model: Claude Sonnet 5; prompt written fresh
Anthropic prompting documentation: ground answers in quotes. Slobodkin et al., ACL 2024: quoting first cuts human verification time
Every record must reference its source text (traceability by design)
Derive the prompt requirements from the hazard model fields (what each field means in the standard) and from what the segments look like
ISO 12100, 3.5 defines harm as "physical injury or damage to health"
Record physical hazards only; duties and legal statements are not hazards
ISO 12100, B.3: hazardous situations are described as tasks
Task field only from segment text or heading path, else empty
ISO 20607, 6.5: warning visibility
Warning marked only on the signal-word line
Kirichenko et al., NeurIPS 2025: models are poor at saying "nothing here"; a careful prompt helps
No-hazard handled by a rule in the prompt
Write the prompt rules based on those requirements and the model guideline
Run them on sample segments, find the errors behind the failures, and update the prompt
Author's choice, labelled as such
One call per segment; two calls tested only if hazards are missed
How to include human review in the framework
Decide the rules of the human expert in the automated framework
Risk evaluation must be done by the expert: its input is not in the manual, and accountability sits with the expert (ISO/TR 14121-2, 5.4)
The expert fills the risk estimation fields
Hazard zone only exists after installation (ISO 12100, 3.11)
The reviewer fills the hazard zone
ISO 20607 requires the warnings, so no warning should be silently lost
A warning with no matching record goes to the reviewer
Only the expert can tell if two duplicate records are one scenario or two (suggestion)
Duplicated hazards across documents are grouped and shown to the reviewer to decide
The reviewer removes invalid and repeated hazards
List the decisions the standards give to the expert (ISO/TR 14121-2, 5.4; hazard zone) (suggestion)
List the cases the pipeline cannot decide alone (duplicates, warnings with no record, invalid records) (suggestion)
Define one reviewer action for each: fill, accept, reject, merge (suggestion)
1 more item...
How to judge if the dedicated pipeline overcomes the LLM agent in the contested issues
Form a pipeline based on the discoveries from the previous investigative questions
Design a prompt for the AI agent based on prompt design instructions to produce a structured output and fulfil the output requirements
Compared condition: Claude as an agent following the pipeline rules, run in Cowork (no API calls); RAG and plain chat dropped
Compare the output from both comparing consistency, coverage, traceability and cost
Use the same AI model for both
Compare the output after human review to filter out repetitions and invalid hazards
Traceability: report the pipeline as 100% by design, test the agent; if it is also 100%, drop traceability as a measure
Assisted hazard identification in reconfigurable manufacturing needs hazard models, and nobody produces them automatically