Digital Uncertainty Laboratory

Tools: Any AI chatbot (ChatGPT, Claude, Gemini), library databases (Scopus, Google Scholar), a shared document for notes
Best for: analysing logical errors, detecting bias

Author: Alina Guzik
Work: Small groups: 3-6 Duration: 46–90 minutes Interaction: High Concentration: High Preparation: High
Digital Uncertainty Laboratory

Why Use It

The method involves the systematic verification of AI-generated content by breaking it down into its constituent elements and assessing the reliability of each of them. Students work with the materials, identify claims requiring verification, compare them against sources, and analyse the coherence of the argumentation. The method develops critical analytical skills, promotes the responsible use of AI tools, and strengthens accountability for the interpretation of information. It demonstrates that correct language and coherent form are not sufficient indicators of quality, and that reliability requires evidence, context, and precision. In practice, this method reproduces the logic of peer review and the information verification processes used in expert environments, where every claim is subject to scrutiny.
In the analytical variant of this method, students are given an AI-generated text and asked to prepare an audit of it: they identify inaccuracies and gaps in the argumentation, supported by references to sources. In the confrontational variant, they formulate prompts in such a way as to encourage the AI to generate potentially problematic content, and then document and classify the resulting errors. In the comparative variant, they compare AI-generated responses with academic literature or expert materials, analysing differences in quality, precision, and style of argumentation.

Practical Example

Course: Research Methodology
Topic: Evaluating the Reliability of Sources and Argumentation in the Age of AI
Teams conduct an in-depth analysis of an AI-generated article addressing a controversial scientific hypothesis. The aim is to identify unverified or fabricated data, logical fallacies, and potential biases in the line of argumentation.

Warm-Up (10 min)

Present the problem: AI claims that [INSERT CONTROVERSIAL CLAIM], citing recognised authorities. Can we trust it? Working in groups, students list three reasons why AI might mislead users, such as limited access to the latest research, a tendency to satisfy user expectations, or cultural oversimplifications.

Provocation and Generation (10 min)
Teams design a prompt intended to encourage AI to generate a text with a high risk of error, for example: Write a scientific article on the advantages of theory X, including expert names, publication dates, and specific statistics. Students then copy the generated output into a shared document.

Substantive Audit – Textual Autopsy (20 min)

Students divide the text into individual claims and subject them to cross-checking.
• Source analysis: Do these authors actually exist? Are the publications indexed in Scopus?
• Logic test: Does the conclusion genuinely follow from the premises presented?
• Language analysis: How much of the text is based on concrete evidence, and how much is merely rhetorical filler?

Hallucination Report (10 min)

Teams prepare an error map, marking false information in red and half-truths or inaccuracies in yellow. Each group presents one concrete piece of evidence confirming an error, for example, identifying an author whose publication could not have appeared at the stated time.

Summary and Reflection (10 min)

Brief discussion: Which AI-generated error was the most difficult to detect, and why?
Class vote for the “Most Brazen AI Hallucination”.

When It Works Best

  • When students already possess basic knowledge of the field and can confidently assess AI-generated content from an expert perspective.
  • Before dissertation writing, to discourage the uncritical use of chatbot-generated sources.

When It Should Be Avoided

  • When working with entirely unfamiliar topics, as students may unknowingly accept AI hallucinations as facts.
  • When reliable databases or source materials are not available for cross-checking.

Challenges

  • Checking everything can be tedious → divide roles within the group: one person checks names, another checks the logic, and a third checks numerical data.

Adjust the Level

Easier → Provide students with a previously generated text containing clearly identifiable errors. Support them with guiding questions.
More Challenging → Students must create a prompt intentionally designed to mislead the AI and then explain the mechanism that caused the chatbot to produce misleading or incorrect information.

Tips

  • Ask students to query AI about the same issue in three different languages, such as Polish, English, and Chinese, and compare the differences in factual content.
  • Reward the group that identifies the most subtle and difficult-to-detect hallucination.
  • If a group fails to identify any errors, point out one hidden error and ask why they overlooked it.

How to Assess

Formative Assessment

Students present successive versions of their prompts and explain what changes they made and why. The teacher comments on their work, asks guiding questions, and provides feedback to help them improve their prompts. The purpose of the assessment is to develop students’ awareness of how the way instructions are formulated influences AI-generated responses.

Summative Assessment
Students submit their final prompt, the response generated by the AI system, and a brief justification of the strategy they adopted. The teacher evaluates the final outcome according to predefined criteria, such as the clarity of the prompt, the relevance of the response, the quality of the justification, and the responsible use of AI.