AI & Compliance

    AI answered 95.8% of dangerous goods questions correctly in one study. So can you trust it?

    DGXprt Team18 September 20268 min read

    It depends on what you asked, how you asked it, and what information the AI models had at the time. New research by the National Cargo Bureau and Durham University tested 13 AI models against the IMDG Code, and the results show why that context matters. The best model outperformed human benchmarks on general knowledge questions, while the same models performed poorly on segregation-related questions, scoring between 18% and 40%.

    One detail in the study that most summaries miss matters: the models were answering from their own training data. The researchers did not give them the IMDG Code to work from.

    DGXprt recently spoke at the AIDGC Annual Conference about AI in dangerous goods compliance, based on our experience building the platform and where the useful part ends. This research backs that view with independent numbers.

    What is DGEval?

    DGEval is a benchmark that measures how accurately AI language models answer dangerous goods compliance questions under the International Maritime Dangerous Goods Code. The National Cargo Bureau's Hazcheck team developed it with Durham University.

    As part of the study, 1,678 questions were put to 13 models from six providers. The questions were categorised into four groups: multiple-choice recall, open-ended questions, structured lookups from the Dangerous Goods List, and identifying the specific IMDG Code provision behind an answer.

    To date, DGEval represents the largest independent test of AI on dangerous goods regulations.

    Where does the AI get its answer from?

    This matters before looking at the results, because it explains most of them. AI tools can arrive at an answer in three different ways, and they are not equally reliable.

    • 1From memory. The model answers from what it absorbed during training: an enormous amount of text collected from the internet. It has no document open in front of it and no way of checking itself. This is what happens when someone types a question into a chatbot with nothing attached.
    • 2By searching the web. The model searches the internet while it answers and uses what it finds. Better, although the quality depends entirely on what happens to be published online and how well the model interprets it.
    • 3From sources it has been given. The model is given the actual documents, the regulation, the standard, the safety data sheet, your own procedures, and is instructed to answer from those alone. For example, ask whether two substances can be stored together, and it answers from the segregation table in front of it and names it. This is how purpose-built compliance software works.

    DGEval tested the first approach and examined the second separately. The study did not include the third approach, so keep that in mind when reading every number that follows.

    What the research found

    Best model, multiple choice

    Top performing AI model on recall questions.

    95.8%

    Experienced human practitioners, multiple choice

    Human benchmark (dangerous goods professionals).

    83.8%

    Classification and documentation questions

    Includes classification, marking, labelling and documentation.

    82% to 85%

    Stowage and segregation questions

    Includes stowage codes, segregation codes and segregation groups.

    around 49% to 50%

    Correctly identifying the IMDG Code provision

    Model identifies the specific clause supporting the answer.

    14.2%

    AI can clearly compete with humans on dangerous goods knowledge. The top model outscored experienced practitioners on recall questions, and that deserves serious attention.

    Now let's examine where AI struggled. On stowage and segregation, models averaged 35 percentage points below their own classification and documentation performance.

    Let's break down the lookup task by field for additional insights and a clearer picture:

    • Stowage Code: 39.8%
    • Segregation Code: 25.5%
    • Segregation Group: 18.4%

    One more result deserves attention. When asked to identify the specific provision in the IMDG Code supporting an answer, the models did so 14.2% of the time. Even if an answer is correct, if it isn't linked back to a particular clause, you have no easy way to verify it. You are back to opening the Code yourself, which is the work you were trying to save.

    "AI wants to be helpful. The problem in compliance is that a plausible answer can look just like the right answer. If you can't trace it back to the source, you don't know which one you have."

    Miki Makuch, Chief Technical Officer, DGXprt, at AIDGC 2026

    The context most coverage leaves out

    Here is the part that changes your perspective on all of the above.

    The researchers ran the test closed-book. According to their report: "No extracts from the published IMDG Code were provided to the evaluated models as source material."

    So, the study measures what these models happened to absorb about the IMDG Code during training. That is a fair and useful test of what happens when somebody types a compliance question into a chatbot. It tells us little about what a purpose-built compliance system can do, because no such system would work that way.

    What happened when the models could reach sources

    The researchers also ran conditions with web search turned on, the second of the three approaches described earlier. The difference was substantial.

    On Dangerous Goods List lookups, accuracy rose by 25 percentage points on average. One smaller model went from 32.2% to 81.9%, a lift of nearly 50 points.

    On multiple-choice questions, web search changed almost nothing, moving results by 0.3 points. That makes sense, since the answer options already narrow the field.

    The pattern is clear. When the task depends on looking something up, access to sources transforms performance. When it depends on recall alone, the model is working from memory, and you get whatever memory it has. That contrast is what matters most here.

    "The question isn't whether AI can help with compliance. It can. The question is what you ask it to do, what information you give it, and what you should never leave it to decide."

    Christine Adeline, Chief Product Officer, DGXprt, at AIDGC 2026

    Why sources alone still leave a gap

    Grounding an AI in the right documents solves much of the problem. It doesn't finish the job, and segregation shows why most clearly, because it depends on more than document access.

    Classification questions usually have a stable answer. A substance carries a UN number, a class, a packing group, and that information is written down in a consistent form in many places.

    Segregation behaves differently. The answer comes from a relationship between two or more things, read from a table, then shaped by context. What else sits in the space, in what quantity, in what packaging, under which jurisdiction, with which controls. Alter one input and the correct answer moves.

    Language models generate text by predicting the next likely word. That suits recall and finding things well. It is a poor fit for relational, table-driven logic where the answer must be exact every single time. The paper's own explanation for the weak segregation results is that these fields are "unlikely to be well-represented in general pre-training corpora."

    A rule of that kind should be calculated against the regulation. Generating it, even from good sources, invites the sort of variability that the numbers above describe.

    What good looks like in practice

    The researchers' conclusion is measured. AI may support compliance work, particularly structured lookups, and human oversight plus verification against authoritative sources remain necessary before it goes anywhere safety-critical, which brings us to the practical implications.

    This is close to what we put to the room at AIDGC, and four habits separate a useful setup from a risky one.

    • Ground it in sources you trust. Decide up front what the authoritative material is, then require the tool to work from that and nothing else. The 25-point lift shows how much this matters.
    • Calculate what should be calculated. Threshold quantities, aggregation, segregation outcomes. These are rules, and rules should be handled by deterministic logic, meaning logic that produces the same answer from the same inputs every time. Ask it twice, get the same answer twice.
    • Insist on citations. With an unaided citation rate of 14.2%, an answer you cannot trace to a clause is an answer you have to verify from scratch.
    • Keep the decision with a person. The higher the consequence, the more oversight the work needs. The rough test we offered at the conference still holds up. If the mistake is an apology, use AI. If the mistake costs money or reputation, use AI with a human reviewing it. If the mistake changes someone's life, a human decides.

    Most dangerous goods work sits in those last two categories.

    Want a practical starting point?

    We have put these habits into a free one-page guide, Getting AI Answers You Can Verify, covering how to frame a question, name your sources, and check the answer before you rely on it.

    Read the free guide

    The architecture that works

    We closed our AIDGC session with a simple picture of how these pieces fit together, and it is worth repeating here because the research maps onto it so directly.

    DGXprt AI architecture showing authoritative information, deterministic logic, AI interpretation and professional judgement

    Where this leaves DGXprt

    That architecture is how DGXprt is built.

    Authoritative regulatory content sits beneath the platform, so the AI never works from memory. Deterministic compliance logic handles the rules, which means segregation outcomes and threshold calculations are computed against the regulation rather than generated. AI works on top of both, helping people find information, understand it, and get clear explanations.

    The accountable professional stays in control of the decision.

    Grounded AI on its own would leave the segregation problem unsolved. Deterministic rules on their own would give you accuracy without accessibility. Together, they make the technology useful and safe, and it is encouraging to see independent research pointing in the same direction, which brings the argument full circle.

    The takeaway

    A general-purpose AI assistant, asked a dangerous goods question with no sources attached, is an unreliable narrator in the areas that matter most operationally. The research makes that clear.

    The same technology, grounded in the right material and paired with rules-based logic, becomes a real asset. The difference lies in how the system is built, and in who stays accountable for the outcome.

    Source: "Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance", National Cargo Bureau Hazcheck team and Durham University. Research summary at natcargo.org/dgeval, full paper at arXiv:2608.21036.

    Note on scope: DGEval tested models against the IMDG Code for maritime transport. Stowage is an IMDG term, covering where goods are placed on a vessel; in Australian land transport and storage the related concepts are loading, storage and segregation. The patterns the study identifies apply broadly to dangerous goods work, though the figures themselves relate to IMDG.

    Stay Informed with DGXprt Quarterly Manifest

    Get expert insights, platform updates, and industry news delivered directly to your inbox every quarter.

    Join industry leaders and compliance professionals staying ahead of dangerous goods regulations.

    Ready to Transform Your Compliance?

    Discover how DGXprt can simplify your dangerous goods management and ensure compliance.