Why this matters on the job
Every Indian consumer app has a day when traffic explodes: a Big Billion Days sale, an IPL final on a streaming app, Tatkal booking at 10 AM, results day on a university portal. Functional tests prove the app works; performance tests prove it keeps working when 50,000 people press the same button at once. Banks, insurers, e-commerce and SaaS product companies in Bengaluru, Pune, Hyderabad and Chennai all hire performance testers, and service companies staff dedicated performance CoEs for their clients.
In interviews you will be asked to explain load vs stress vs soak, why the 95th percentile matters more than the average, and to calculate how many virtual users you need for a target throughput. This lesson gives you that vocabulary and, more importantly, lets you measure real numbers with a tiny script before we touch JMeter.
Concepts
Types of performance tests
| Test type | Question it answers | Typical shape |
|---|---|---|
| Load test | Does the system meet NFRs at the expected peak load? | Ramp up to peak, hold 30-60 min, ramp down |
| Stress test | Where does it break, and how does it fail? | Keep increasing load past peak until errors/latency explode |
| Spike test | Can it absorb a sudden burst (sale starts, push notification)? | Normal load, jump to 5-10x in seconds, drop back |
| Soak / endurance | Are there memory leaks, connection leaks, log/disk growth? | Moderate load for 4-12+ hours |
| Scalability test | Does adding pods/CPU increase capacity proportionally? | Same test on 1, 2, 4 nodes; compare max throughput |
| Volume test | Does it slow down with large data (10 crore rows)? | Normal load against a production-sized database |
| Smoke perf test | Is the script and environment sane before a big run? | 1-5 users, few minutes |
The KPIs you report
- Response time - time from sending the request to receiving the last byte. Report percentiles (p50/median, p90, p95, p99), not just the average. p95 = 2 s means 95% of requests finished in 2 s or less.
- Latency / TTFB - time to the first byte; shows server think time separately from download time.
- Throughput - completed requests or business transactions per second (TPS/RPS). Also network throughput in KB/s.
- Error % - failed requests (HTTP errors, timeouts, failed assertions) divided by total.
- Concurrency - active virtual users (VUs) or in-flight requests.
- Resource utilisation - CPU, memory, GC, DB connections, thread pools on the servers.
- Apdex - (satisfied + tolerating/2) / total, a 0-1 user-satisfaction score used in JMeter dashboards and APM tools.
Little's law: the formula interviewers love
For a stable system: N = X × (R + Z) where N = concurrent users, X = throughput (req/s), R = response time (s), Z = think time (s). Example: the business wants 100 searches/s, a search takes 0.5 s and users think for 4.5 s between actions. You need N = 100 × (0.5 + 4.5) = 500 virtual users. With zero think time, 50 users would produce the same 100 req/s - which is why removing think time makes a test unrealistically aggressive.
Writing a good NFR
“The site should be fast” is not testable. A good non-functional requirement is SMART: it names the transaction, the load, the percentile, the limit, the error budget, the duration and the environment. Ask the product owner or derive it from production analytics (peak hour from Google Analytics, access logs or APM).
NFR-PERF-01 Find Flights (POST /reserve.php)
Load : 200 concurrent users, 5 s think time, 30 min steady state
Response : p95 <= 2.0 s, p99 <= 4.0 s (measured at the load injector)
Throughput: >= 35 transactions/s sustained
Errors : < 1 % (HTTP 4xx/5xx + failed assertions)
Resources : app server CPU < 70 %, memory with no upward trend
Data : 10,000 bookings in DB before test; 50 unique routes in test dataHands-on Lab: Measure percentiles, throughput and Little's law with real traffic
- Install Python 3.8+ (
python3 --version). No libraries are needed. - Break one request into phases with curl (Git Bash on Windows works too). Run it 3 times and note how DNS/TLS drop after the first call because of caching:
Expected: something likecurl -o /dev/null -s -w "dns=%{time_namelookup}s connect=%{time_connect}s tls=%{time_appconnect}s ttfb=%{time_starttransfer}s total=%{time_total}s code=%{http_code}\n" https://blazedemo.com/dns=0.01s connect=0.2s tls=0.45s ttfb=0.7s total=0.72s code=200. TTFB minus TLS is roughly server processing time. - See why averages lie. Open a Python shell and run:
The average (440 ms) describes nobody: 95 users saw 200 ms and 5 users waited 5 s. p99 exposes them.>>> import statistics, math >>> times = [200] * 95 + [5000] * 5 # 100 requests, 5 very slow ones >>> statistics.mean(times) 440 >>> s = sorted(times) >>> s[math.ceil(0.95 * 100) - 1], s[math.ceil(0.99 * 100) - 1] (200, 5000) - Save the following as
perf_basics.pyand runpython3 perf_basics.py:# perf_basics.py -- run: python3 perf_basics.py (Python 3.8+, no extra packages) import math import statistics import time import urllib.request from concurrent.futures import ThreadPoolExecutor URL = "https://blazedemo.com/" REQUESTS = 30 # keep it small: this is a shared public demo site WORKERS = 5 # 5 concurrent "users" with zero think time def hit(_): start = time.perf_counter() try: with urllib.request.urlopen(URL, timeout=15) as resp: resp.read() ok = resp.status == 200 except Exception: ok = False return (time.perf_counter() - start) * 1000, ok def pct(values, p): """Nearest-rank percentile (same idea JMeter/k6 reports use).""" s = sorted(values) k = max(0, math.ceil(p / 100 * len(s)) - 1) return s[k] if __name__ == "__main__": t0 = time.perf_counter() with ThreadPoolExecutor(max_workers=WORKERS) as pool: results = list(pool.map(hit, range(REQUESTS))) wall = time.perf_counter() - t0 times = [r[0] for r in results] errors = sum(1 for r in results if not r[1]) avg_s = statistics.mean(times) / 1000 throughput = len(times) / wall print(f"samples={len(times)} errors={errors} ({errors / len(times):.1%})") print(f"avg={statistics.mean(times):.0f} ms median={pct(times, 50):.0f} ms") print(f"p90={pct(times, 90):.0f} ms p95={pct(times, 95):.0f} ms " f"p99={pct(times, 99):.0f} ms max={max(times):.0f} ms") print(f"throughput={throughput:.2f} req/s over {wall:.1f} s") # Little's law: N = X * R -> predicted concurrency should be close to WORKERS print(f"Little's law check: X * R = {throughput * avg_s:.2f} (workers = {WORKERS})") - Record the output in a table: avg, median, p90, p95, p99, max, error %, throughput. Expected order: median ≤ p90 ≤ p95 ≤ p99 ≤ max, and the Little's law line should print a value close to 5 (the number of workers), because with no think time N = X × R.
- Change
WORKERSto 1 and then 10 (keep REQUESTS at 30). Observe: throughput rises with workers while the server is not saturated; if response time also rises sharply, you are seeing queueing. - Using Little's law, calculate the VUs needed for 40 TPS on blazedemo if R = 0.8 s and Z = 4 s (answer: 40 × 4.8 = 192 VUs). Write the calculation in your notes.
- Write three NFRs for blazedemo (Home, Find Flights, Purchase) in the format shown above. You will reuse them in the capstone.
Ethics first: blazedemo.com and test.k6.io exist for learning, but they are shared. Keep public-site tests small (tens of users, minutes not hours). Heavy stress tests belong on your own environment - we will use a local Docker app in Lesson 6.
Common mistakes
- Reporting only the average response time - always give p90/p95/p99 and the max.
- Testing without agreed NFRs, then arguing about whether 3 s is “good”.
- Removing think time to “generate more load” - it changes the user mix and inflates throughput unrealistically.
- Mixing up response time measured at the client (includes network) with server time from APM.
- Running a load test against production or a shared demo site without written permission.
- Calling a 10-minute run a soak test - leaks need hours to show up.
Real-world assignment
Your manager forwards a mail: “Diwali sale expected 3x normal traffic. Normal peak = 1,200 orders/hour, average 6 page views per order, median page response 0.9 s, user think time 8 s. Tell us how many VUs to simulate and what NFRs we should sign off.” Produce a one-page note with: peak TPS (pages/s), required VUs via Little's law, a test-type plan (smoke, load, stress, spike, 4-hour soak) and five SMART NFRs. Hint: 3 × 1,200 × 6 / 3600 = 6 pages/s; N = 6 × (0.9 + 8) ≈ 54 VUs at the 3x peak, so plan the stress test to go well beyond that.
Key takeaways
- Load = expected peak, stress = beyond peak, spike = sudden burst, soak = long duration, scalability = more hardware, more capacity?
- Percentiles (p95/p99) describe user experience; averages hide outliers.
- Little's law N = X × (R + Z) converts business throughput into virtual users.
- Throughput, response time, error % and server utilisation must be read together.
- Every test needs SMART NFRs and a target environment you are allowed to load.