Case File 97-JJI All posts
Case 97-JJIBlog

The AI Reached the Conclusion, but You Answer for It

Over my break, I developed a new lecture for next semester about the use and ethics of AI in digital forensic investigations. I'll open this lecture with a show of hands. "Who in the room has used an AI tool for work or school in the last 30 days?" Most hands will go up. The second question follows: "Who checked what that tool kept, or where the data went?" The hands will come down quickly, I'm sure.

That gap—between how fluently we use these tools and how little we examine them—is the whole subject. AI capability moves fast. Regulation moves slowly. Professional standards move slower still. Every problem I raise in this material lives somewhere in the space between those three speeds, and when capability outruns the rules, someone is still standing at the workbench who has to decide what to do. The absence of a rule is not the absence of a duty.

AI shows up in this profession wearing three different hats. Sometimes it's a tool you run to do the examination. Sometimes it's an actor that produces something carrying evidentiary weight. Sometimes it's the subject—the platform or the data you're investigating. The ethics change with the role, and tracking which one is in play is the first discipline, because most of the trouble starts when an examiner reaches for the tool without registering that it's also become the actor.

Start with the tool. The gains are real, and that's exactly what makes the rest of it matter. AI is good at triage at scale, malware behavioral classification, anomaly detection across network traffic, and language analysis over document sets no human team could read in a year. The catch arrives the moment the tool stops matching against material a human already identified and starts judging content it has never seen. A hash hit against a known set rests on a prior human decision, so your judgment can be light. A classifier that reports "92% likely contraband" has handed you a probability estimate and nothing more. A probability threshold is not an ethical framework, and the call—report, testimony, and charge—is still yours.

This is where explainability stops being academic. If you can't explain how the tool reached its result, you can't honestly call that result your expert analysis. The algorithm is not a co-author. You signed the report, you hold the certification, and you answer for it on the stand. Point the Daubert conditions at the tool the way a defense attorney will—reliability, testability, a known error rate, and peer review—and many AI forensic tools can't satisfy a single one. State v. Loomis showed the shape of the fight: a risk score used at sentencing with its methodology held as a trade secret, a defendant arguing he couldn't challenge what he wasn't allowed to see, a state court that let it stand with cautions, and a U.S. Supreme Court that declined to take the case. Same black box, same objection, now pointed at your evidence.

The fairness problem runs deeper than any single tool, because a model learns from history, and history carries every distortion of the world that produced it. The bias is built into the architecture. No patch is coming for it. Robert Williams spent a night in a Detroit cell in 2020 because a facial-recognition system returned a false match and the officers ran with it—during interrogation, he was told "the computer says it's you." NIST's testing has found demographic differentials across most algorithms it has examined, with false-positive rates running higher for some groups and the gap varying enormously by tool. Using a tool blind to its own error profile, on the face in front of you, is a choice with consequences. The same loop runs through predictive policing, where biased enforcement data trains the model, the model sends officers back to the same blocks, those patrols generate more of the same data, and the loop closes from the inside.

The habit the professional literature has barely touched is the LLM used off the books, mid-case. Examiners paste scripts, chat logs, emails, and metadata into ChatGPT, Claude, Gemini, Grok, or Copilot because it's fast and often useful. Commercial models keep what you give them, and the terms change without notice. Feeding live evidence to a public model raises real questions about whether you've broken chain of custody, waived privilege in a commercial matter, or violated a protective order in civil litigation. The doctrine is unsettled. The habit shouldn't wait for it—document the AI use now, before a court tells you what you should have done.

The hardest version of this is AI-generated child sexual abuse material, where the law splits on one question: whether a real child is involved. Material tied to an identifiable child, including face-swaps and morphs onto a real child, falls under the traditional statutes. Wholly synthetic material runs into the child-obscenity statute and the protection the Supreme Court extended to non-obscene virtual content in Free Speech Coalition. United States v. Anderegg is the live test, with a private-possession count dismissed in early 2025 on Stanley v. Georgia grounds while production and distribution counts proceed—a district-court ruling, not the last word. At discovery you usually can't tell origin or obscenity, so the defensible default is to preserve, decline to distribute, and report through proper channels. The volume tells its own story: NCMEC's AI-CSAM reports went from roughly 6,800 in 2024 to more than 440,000 in the first half of 2025. At that scale, automated triage stops being a convenience and becomes occupational health for the people doing the work.

Flip the script and the platform becomes the target. What an AI system holds—conversation history, prompts, outputs, account metadata, payment records, and usage logs—is potentially probative, if it was kept, and for how long. Retention varies by provider, by tier, and by the week. Whether an LLM transcript even counts as a "stored communication" under the Stored Communications Act is unsettled. Underneath all of it sits the attribution problem: when an autonomous system causes harm, the subject of your investigation could be the developer, the deployer, the operator, or the person who prompted it. ECPA dates to 1986. Fourth Amendment doctrine on AI-held records is barely sketched. These are the working conditions you graduate into.

This brings the whole thing back to the signature. Every ethics code in this field says stand behind what you sign, and the algorithm shares none of it—not your liability, not your oath, and not your need to keep your job. When an AI finding turns out to be wrong, the weight is supposed to distribute across examiner, agency, and vendor. Vendors disclaim by contract. Agencies point to procurement. Weight that won't distribute cleanly concentrates, and it concentrates on the name at the bottom of the report. The institution adopts the tool for throughput and cost. The individual carries it case by case and signs personally.

This is the part the field keeps deferring, and it can't defer much longer. The standards for AI in digital forensics will be written in the next few years, and everyone graduating into this work will operate under them regardless. Professions that fail to govern themselves get governed, and rules drafted by non-practitioners tend to be blunt instruments. The window to shape this is open right now. It won't stay open.

For practitioners reading this: what has your organization actually put in writing about AI use—tool validation, documentation of which models touched an analysis, and where live evidence is permitted to go—and where are the gaps you already know exist? For those in adjacent fields, where a finding still gets attached to a person's name, how is your profession handling the same gap between what the tools can do and what the rules have caught up to?


This post is the eleventh and final in a series based on my course—DFOR 671: Topics of Ethics and Law in Computer Forensics, that I have taught at George Mason University for the past fifteen years. We started with the question every new class gets on day one—is digital forensics an art or a science—and worked through professional honesty, bias, the human dynamics of a search, Fourth Amendment scope, expert testimony, documentation, OSINT, commercial practice, and the mental model gap between law enforcement and intelligence work, before arriving here, where the tools themselves are generating the hardest questions. The ethical terrain underneath this field gets harder as the technology gets more capable. I hope you've enjoyed reading these as much as I've enjoyed writing them.

First published on LinkedIn.