Why this matters on the job
Open almost any QA job description on Naukri or LinkedIn today and you will find lines such as "exposure to GenAI tools", "prompt engineering for test design" or "experience with GitHub Copilot". Large service companies have rolled out AI coding assistants to thousands of engineers, and product companies expect testers to use AI to move faster. At the same time, managers have seen AI-written test cases that test features which do not exist, and automation scripts that pass while checking nothing.
The tester who wins is not the one who uses AI the most, but the one who knows where AI is reliable, where it is dangerous, and how to verify its output. This lesson gives you that mental model and a small experiment you will repeat throughout the course.
Concepts
AI, ML, deep learning, GenAI and LLMs in one table
| Term | What it means | Testing example |
|---|---|---|
| Artificial Intelligence (AI) | Any technique that makes software behave 'intelligently' | A tool that picks which regression tests to run |
| Machine Learning (ML) | Models that learn patterns from data instead of hand-written rules | Flaky-test prediction from past CI runs |
| Deep Learning | ML using large neural networks | Visual comparison that ignores anti-aliasing noise |
| Generative AI (GenAI) | Models that generate new text, code or images | Drafting test scenarios from a user story |
| Large Language Model (LLM) | A GenAI model trained on huge text/code corpora to predict the next token | ChatGPT, Claude, Gemini, the model behind GitHub Copilot |
How an LLM actually produces an answer
- Tokens: text is split into small pieces (roughly 3-4 characters of English each). Pricing and limits are counted in tokens.
- Next-token prediction: the model repeatedly predicts the most plausible next token. It is optimised for plausible, not for true.
- Context window: the amount of text (prompt + documents + answer) the model can 'see' at once. Anything outside it does not exist for the model.
- Temperature / sampling: controls randomness. Even at low temperature, most hosted chat tools can return different answers to the same prompt.
- Knowledge cutoff: the model's training data stops at a date; newer library versions or APIs may be unknown unless the tool searches the web.
Where AI helps across the STLC
| STLC phase | Where AI helps | What the human must still do |
|---|---|---|
| Requirement analysis | Summarise stories, list ambiguities and open questions | Confirm questions with BA/PO; decide scope |
| Test planning | Draft plan sections, risk lists, estimation checklists | Own risk priority, dates, resourcing |
| Test design | Scenarios, edge cases, BVA/EP tables, Gherkin, synthetic data | Trace to acceptance criteria, remove invented behaviour |
| Automation | Generate/refactor Playwright, Selenium, API code | Review, run, make it fail on purpose, keep standards |
| Execution and defects | Structure bug reports, summarise logs, cluster failures | Reproduce, attach real evidence, decide severity |
| Closure and reporting | Narrative for summary reports, release notes | Verify every number from the source tool |
Limitations you must design around
| Limitation | What it looks like in QA work | Counter-measure |
|---|---|---|
| Hallucination | A Selenium method or Postman API that does not exist; a requirement nobody wrote | Check against official docs and the real app; run the code |
| Non-determinism | Same prompt gives 12 scenarios today, 17 tomorrow | Save prompts and outputs; review the final artefact, not the chat |
| Outdated knowledge | Selenium 3 style waits, deprecated Cypress commands | State versions in the prompt; read release notes |
| Confident tone | Wrong answers sound as sure as right ones | Ask for sources/assumptions; verify independently |
| Privacy risk | Pasting client logs, customer data or source code into a public chatbot | Follow company AI policy; mask data; use approved tools only |
| Weak arithmetic and counting | Wrong pass percentages in a summary report | Compute numbers with code or the test tool |
Two different skills: testing with AI (using AI to help you test normal software, lessons 2-7 and 9-10) and testing AI (testing products that contain an LLM, lesson 8). Interviewers increasingly ask about both.
Hands-on Lab: Measure hallucination and non-determinism yourself
You need any chat assistant you are allowed to use (ChatGPT, Claude, Gemini or Copilot Chat), Python 3.10+ and a browser. Use only public information - we test the public demo shop saucedemo.com.
- Create a folder
ai-lab-01. Open a new chat and send exactly this prompt:You are a QA engineer. List 8 test cases for the login page of https://www.saucedemo.com. For each give: ID, title, test data, expected result. Use a plain numbered list. - Copy the answer into
run1.txt. Open two more new chats, send the identical prompt and saverun2.txtandrun3.txt. - Save this script as
compare_runs.pyin the same folder and runpython compare_runs.py:import difflib import itertools import pathlib import sys files = sorted(pathlib.Path(".").glob("run*.txt")) if len(files) < 2: sys.exit("Save at least two AI responses as run1.txt, run2.txt ...") texts = {f.name: f.read_text(encoding="utf-8") for f in files} for a, b in itertools.combinations(texts, 2): ratio = difflib.SequenceMatcher(None, texts[a], texts[b]).ratio() print(f"{a} vs {b}: {ratio:.0%} similar") line_sets = [ {line.strip().lower() for line in t.splitlines() if line.strip()} for t in texts.values() ] common = set.intersection(*line_sets) print(f"Lines identical in ALL runs: {len(common)}") - Expected result: similarity is usually well below 100% and very few lines are identical across all runs. Write one sentence in
findings.md: "The same prompt gave different test sets; I must review the final artefact, not trust a single run." - Fact check: open saucedemo.com. The login page itself lists the accepted usernames (standard_user, locked_out_user, problem_user, performance_glitch_user, error_user, visual_user) and the password
secret_sauce. Compare with the test data the AI used. Mark each AI test case as Correct, Wrong data or Invented behaviour (for example "account locks after 3 failed attempts" - saucedemo has no such rule). - Hallucination probe: in a new chat ask: "In Selenium 4 Java, how do I use driver.findElementByText() to click the Login button?" No such method exists in Selenium WebDriver; Selenium 4 locates elements with
driver.findElement(By...). Note whether the assistant corrects you or invents usage. Verify on the official Selenium finders page. - Fill this verification log in
findings.mdand keep it - you will extend it in later lessons:| Prompt | Tool + date | Claim checked | Source used to verify | Verdict | |--------|-------------|---------------|-----------------------|---------| | Login TCs run1 | <tool>, <date> | locked_out_user error text | saucedemo.com live | Correct | | findElementByText | <tool>, <date> | method exists | selenium.dev docs | Hallucinated |
Common mistakes
- Treating the first AI answer as final. One run is a sample, not a specification.
- Pasting client requirements, production logs or customer data into a personal chatbot account. This can breach NDAs and India's Digital Personal Data Protection (DPDP) Act, 2023.
- Asking AI "is this correct?" about its own answer and accepting "yes". Verify with docs, the app, or by running the code.
- Assuming AI knows your app. It only knows what is in the prompt and its training data.
- Using AI output as test evidence. Evidence is what the system actually did: screenshots, logs, reports.
Real-world assignment
Your QA lead says: "Management wants us to start using AI. Write a one-page note for our project." Produce ai-usage-note.md with: (1) five concrete tasks in your project's STLC where AI will be used, (2) the approved tool for each (as per your company policy), (3) five rules, for example never paste production data, every AI artefact is peer reviewed, numbers come from tools, not from AI, and (4) how you will measure benefit (time saved per story, defects found by AI-suggested scenarios). Keep it vendor-neutral so it survives a tool change.
Key takeaways
- An LLM predicts plausible text; plausibility is not correctness.
- AI can assist every STLC phase, but a human owns scope, evidence, severity and sign-off.
- Hallucination, non-determinism and outdated knowledge are normal behaviour - design your process around them.
- Keep a verification log: prompt, tool, date, what you checked and against which source.
- Data privacy comes first: only approved tools, only masked or public data.