Task
I’m raising a labeling question, not a case.
Last month, a member posted a case. They said an AI tool gave them a wrong answer about a local regulation. They included the claim, the date, the tool, and a screenshot of the output. The evidence looked solid. I believed them.
Then I tried to reproduce it. I used the same tool. I asked the same question. I got a different answer—and this time, the answer was correct.
I tried three more times. Every time, the answer was correct. I couldn’t reproduce the error.
I went back to the original post. The member had deleted their account. The screenshot was still there, but the exact prompt wasn’t included. The model version wasn’t specified. The date was approximate.
So here we are. The case is probably real. The error probably happened. But we can’t reproduce it.
What label should we use?
Tool and Date
This isn’t a tool test. This is a forum governance question.
Context: A prior case in Case Files that could not be reproduced.
Date of the original case: Approximately August 2026.
Date of my reproduction attempt: September 10, 2026.
Tool used in the original case: A general-purpose AI assistant (free tier).
Tool I used to attempt reproduction: Same tool, same tier.
Original Input
The original post is now deleted. But here’s what I remember.
The member asked: “Does [my city] require a permit for a small home business?”
The AI answered: “No, home businesses are exempt from permits in [my city].”
The member later checked the city’s website and found that permits are required for home businesses with more than one employee. They posted the case with the claim, the date, and a screenshot.
Claim Under Review
The claim under review is: “The AI’s answer was wrong, and the error is real.”
I believe the claim. I’ve seen similar errors. I’ve made similar mistakes. But I can’t prove it. When I tried to reproduce it, the AI gave a correct answer.
This raises the labeling question.
Evidence
Here’s what we have.
What supports the claim:
The original member’s screenshot.
The original member’s description of the city’s website.
My own experience with similar errors.
What doesn’t support the claim:
I can’t reproduce the error.
The original prompt isn’t available.
The model version isn’t specified.
The original member is gone.
What’s missing:
A reproducible prompt.
A model version.
A second witness.
A timestamp that matches a known model release.
Finding
Unresolved — and possibly unreproducible.
The current labels are:
Confirmed: The error is real and verified.
Corrected: The error was found and fixed.
Partly correct: Some claims were wrong, others were right.
Unresolved: We can’t tell whether the error is real.
Unresolved is close, but it doesn’t quite fit. Unresolved implies we tried and couldn’t determine the answer. This case is different. We think the answer is probably real. We just can’t reproduce it.
Maybe we need a new label: Unreproducible.

Risk
Here’s why this matters.
Without a label, cases get lost. If a case can’t be resolved, it falls off the front page. No one learns from it. No one tracks the pattern.
Without a label, patterns disappear. If multiple members report the same unreproducible error, we can’t see the pattern if each case is labeled differently.
Without a label, trust erodes. If unresolved cases aren’t clearly marked, readers may assume they were resolved. Or they may assume they were dismissed. Neither is accurate.
Without a label, members stop reporting. If you report an error and it can’t be reproduced, and there’s no place for it, you might not report the next one.
Lesson
Here’s what I’ve learned from this situation.
Reproducibility is not the same as truth. A real error can be unreproducible. A false claim can be reproducible. The two are different.
Model updates break reproducibility. If a model changes, the same prompt can produce a different answer. That doesn’t mean the original error wasn’t real.
Missing prompts break reproducibility. Without the exact prompt, we can’t reproduce the exact output.
Missing versions break reproducibility. Without the model version, we can’t know what the model was capable of at that time.
We need a place for these cases. Not in Case Files. Not in Error Patterns. Somewhere in The Commons, where governance and unresolved questions live.
What I’m Asking the Community
How should the forum label the case?
Here are some options I’ve considered.
Option 1: Use “Unresolved” for everything. Keep it simple. If we can’t resolve it, it’s unresolved. No new labels.
Option 2: Add “Unreproducible” as a sub-label. Keep “Unresolved” for cases we haven’t resolved yet. Add “Unreproducible” for cases we’ve tried and can’t reproduce.
Option 3: Add “Stale” as a label. If a case is older than 90 days and can’t be reproduced, mark it stale. It stays visible but flagged.
Option 4: Add “Reported” as a label. If a member reports an error and we can’t verify it, mark it “Reported.” It’s a signal, not a verdict.
Option 5: Something else. Maybe there’s a better word. Maybe there’s a better system.
I’m also curious about the process. Should we attempt reproduction before labeling? Should we require a prompt and version? Should we allow anonymous reports?
And one more question: what should happen to unreproducible cases? Should they stay visible? Should they be archived? Should they be listed somewhere for pattern-spotting?
I’d like to hear from the community. This is a governance question, not a technical one. I don’t have a strong opinion. I just want a system that works.
If you have a preference, please share it. If you’ve seen other communities handle this well, tell me. If you think we don’t need a new label, tell me that too.
Comments
No comments yet — be the first to share a thought.