Lesson 1 of 10 · Lesson + Lab · 60 min · Free preview

AI, ML and LLMs for Testers — What AI Can and Cannot Do in the STLC

Understand how LLMs work, where AI genuinely helps across the STLC, and measure hallucination and non-determinism yourself with a hands-on experiment.

Why this matters on the job

Open almost any QA job description on Naukri or LinkedIn today and you will find lines such as "exposure to GenAI tools", "prompt engineering for test design" or "experience with GitHub Copilot". Large service companies have rolled out AI coding assistants to thousands of engineers, and product companies expect testers to use AI to move faster. At the same time, managers have seen AI-written test cases that test features which do not exist, and automation scripts that pass while checking nothing.

The tester who wins is not the one who uses AI the most, but the one who knows where AI is reliable, where it is dangerous, and how to verify its output. This lesson gives you that mental model and a small experiment you will repeat throughout the course.

Concepts

AI, ML, deep learning, GenAI and LLMs in one table

TermWhat it meansTesting example
Artificial Intelligence (AI)Any technique that makes software behave 'intelligently'A tool that picks which regression tests to run
Machine Learning (ML)Models that learn patterns from data instead of hand-written rulesFlaky-test prediction from past CI runs
Deep LearningML using large neural networksVisual comparison that ignores anti-aliasing noise
Generative AI (GenAI)Models that generate new text, code or imagesDrafting test scenarios from a user story
Large Language Model (LLM)A GenAI model trained on huge text/code corpora to predict the next tokenChatGPT, Claude, Gemini, the model behind GitHub Copilot

How an LLM actually produces an answer

  • Tokens: text is split into small pieces (roughly 3-4 characters of English each). Pricing and limits are counted in tokens.
  • Next-token prediction: the model repeatedly predicts the most plausible next token. It is optimised for plausible, not for true.
  • Context window: the amount of text (prompt + documents + answer) the model can 'see' at once. Anything outside it does not exist for the model.
  • Temperature / sampling: controls randomness. Even at low temperature, most hosted chat tools can return different answers to the same prompt.
  • Knowledge cutoff: the model's training data stops at a date; newer library versions or APIs may be unknown unless the tool searches the web.

Where AI helps across the STLC

STLC phaseWhere AI helpsWhat the human must still do
Requirement analysisSummarise stories, list ambiguities and open questionsConfirm questions with BA/PO; decide scope
Test planningDraft plan sections, risk lists, estimation checklistsOwn risk priority, dates, resourcing
Test designScenarios, edge cases, BVA/EP tables, Gherkin, synthetic dataTrace to acceptance criteria, remove invented behaviour
AutomationGenerate/refactor Playwright, Selenium, API codeReview, run, make it fail on purpose, keep standards
Execution and defectsStructure bug reports, summarise logs, cluster failuresReproduce, attach real evidence, decide severity
Closure and reportingNarrative for summary reports, release notesVerify every number from the source tool

Limitations you must design around

LimitationWhat it looks like in QA workCounter-measure
HallucinationA Selenium method or Postman API that does not exist; a requirement nobody wroteCheck against official docs and the real app; run the code
Non-determinismSame prompt gives 12 scenarios today, 17 tomorrowSave prompts and outputs; review the final artefact, not the chat
Outdated knowledgeSelenium 3 style waits, deprecated Cypress commandsState versions in the prompt; read release notes
Confident toneWrong answers sound as sure as right onesAsk for sources/assumptions; verify independently
Privacy riskPasting client logs, customer data or source code into a public chatbotFollow company AI policy; mask data; use approved tools only
Weak arithmetic and countingWrong pass percentages in a summary reportCompute numbers with code or the test tool

Two different skills: testing with AI (using AI to help you test normal software, lessons 2-7 and 9-10) and testing AI (testing products that contain an LLM, lesson 8). Interviewers increasingly ask about both.

Hands-on Lab: Measure hallucination and non-determinism yourself

You need any chat assistant you are allowed to use (ChatGPT, Claude, Gemini or Copilot Chat), Python 3.10+ and a browser. Use only public information - we test the public demo shop saucedemo.com.

  1. Create a folder ai-lab-01. Open a new chat and send exactly this prompt:
    You are a QA engineer. List 8 test cases for the login page of https://www.saucedemo.com.
    For each give: ID, title, test data, expected result. Use a plain numbered list.
  2. Copy the answer into run1.txt. Open two more new chats, send the identical prompt and save run2.txt and run3.txt.
  3. Save this script as compare_runs.py in the same folder and run python compare_runs.py:
    import difflib
    import itertools
    import pathlib
    import sys
    
    files = sorted(pathlib.Path(".").glob("run*.txt"))
    if len(files) < 2:
        sys.exit("Save at least two AI responses as run1.txt, run2.txt ...")
    
    texts = {f.name: f.read_text(encoding="utf-8") for f in files}
    for a, b in itertools.combinations(texts, 2):
        ratio = difflib.SequenceMatcher(None, texts[a], texts[b]).ratio()
        print(f"{a} vs {b}: {ratio:.0%} similar")
    
    line_sets = [
        {line.strip().lower() for line in t.splitlines() if line.strip()}
        for t in texts.values()
    ]
    common = set.intersection(*line_sets)
    print(f"Lines identical in ALL runs: {len(common)}")
  4. Expected result: similarity is usually well below 100% and very few lines are identical across all runs. Write one sentence in findings.md: "The same prompt gave different test sets; I must review the final artefact, not trust a single run."
  5. Fact check: open saucedemo.com. The login page itself lists the accepted usernames (standard_user, locked_out_user, problem_user, performance_glitch_user, error_user, visual_user) and the password secret_sauce. Compare with the test data the AI used. Mark each AI test case as Correct, Wrong data or Invented behaviour (for example "account locks after 3 failed attempts" - saucedemo has no such rule).
  6. Hallucination probe: in a new chat ask: "In Selenium 4 Java, how do I use driver.findElementByText() to click the Login button?" No such method exists in Selenium WebDriver; Selenium 4 locates elements with driver.findElement(By...). Note whether the assistant corrects you or invents usage. Verify on the official Selenium finders page.
  7. Fill this verification log in findings.md and keep it - you will extend it in later lessons:
    | Prompt | Tool + date | Claim checked | Source used to verify | Verdict |
    |--------|-------------|---------------|-----------------------|---------|
    | Login TCs run1 | <tool>, <date> | locked_out_user error text | saucedemo.com live | Correct |
    | findElementByText | <tool>, <date> | method exists | selenium.dev docs | Hallucinated |

Common mistakes

  • Treating the first AI answer as final. One run is a sample, not a specification.
  • Pasting client requirements, production logs or customer data into a personal chatbot account. This can breach NDAs and India's Digital Personal Data Protection (DPDP) Act, 2023.
  • Asking AI "is this correct?" about its own answer and accepting "yes". Verify with docs, the app, or by running the code.
  • Assuming AI knows your app. It only knows what is in the prompt and its training data.
  • Using AI output as test evidence. Evidence is what the system actually did: screenshots, logs, reports.

Real-world assignment

Your QA lead says: "Management wants us to start using AI. Write a one-page note for our project." Produce ai-usage-note.md with: (1) five concrete tasks in your project's STLC where AI will be used, (2) the approved tool for each (as per your company policy), (3) five rules, for example never paste production data, every AI artefact is peer reviewed, numbers come from tools, not from AI, and (4) how you will measure benefit (time saved per story, defects found by AI-suggested scenarios). Keep it vendor-neutral so it survives a tool change.

Key takeaways

  • An LLM predicts plausible text; plausibility is not correctness.
  • AI can assist every STLC phase, but a human owns scope, evidence, severity and sign-off.
  • Hallucination, non-determinism and outdated knowledge are normal behaviour - design your process around them.
  • Keep a verification log: prompt, tool, date, what you checked and against which source.
  • Data privacy comes first: only approved tools, only masked or public data.

🧪 Practical checklist

Do each task yourself and tick it off. All tasks are required to complete this lesson.

🎤 Interview questions

How would you explain an LLM to a non-technical manager?

It is a model trained on huge amounts of text that predicts the next most likely word. That makes it great at drafting and summarising, but it can sound confident while being wrong, so its output needs verification.

Where in the STLC have you used AI, and how did you control quality?

Mainly in test design and automation drafting. Every AI artefact was traced to acceptance criteria, run against the real application, and peer reviewed before being added to the suite.

What is hallucination and how does it affect testers?

Hallucination is when the model produces plausible but false content, such as a non-existent API or an invented requirement. Testers counter it by verifying against official documentation and the actual system.

Why can the same prompt give different results, and what do you do about it?

Generation involves sampling, so outputs vary run to run. I save prompts and outputs and review the final artefact rather than relying on any single run.

📝 Quiz, progress tracking & certificate

You're reading a free preview. Premium members tick off labs, take the quiz, unlock all 10 lessons and earn a verifiable certificate.

💎 Unlock with Premium Premium login
✓ You're subscribed! Job alerts arrive daily at 9 AM.
Scroll to Top