Lesson 1 of 10 · Lesson + Lab · 60 min · Free preview

Performance Testing Fundamentals: Test Types, KPIs & NFRs

Learn load, stress, spike, soak and scalability testing, the KPIs that matter (percentiles, throughput, error %), Little's law and how to write testable NFRs - then measure real numbers with a small Python script.

Why this matters on the job

Every Indian consumer app has a day when traffic explodes: a Big Billion Days sale, an IPL final on a streaming app, Tatkal booking at 10 AM, results day on a university portal. Functional tests prove the app works; performance tests prove it keeps working when 50,000 people press the same button at once. Banks, insurers, e-commerce and SaaS product companies in Bengaluru, Pune, Hyderabad and Chennai all hire performance testers, and service companies staff dedicated performance CoEs for their clients.

In interviews you will be asked to explain load vs stress vs soak, why the 95th percentile matters more than the average, and to calculate how many virtual users you need for a target throughput. This lesson gives you that vocabulary and, more importantly, lets you measure real numbers with a tiny script before we touch JMeter.

Concepts

Types of performance tests

Test typeQuestion it answersTypical shape
Load testDoes the system meet NFRs at the expected peak load?Ramp up to peak, hold 30-60 min, ramp down
Stress testWhere does it break, and how does it fail?Keep increasing load past peak until errors/latency explode
Spike testCan it absorb a sudden burst (sale starts, push notification)?Normal load, jump to 5-10x in seconds, drop back
Soak / enduranceAre there memory leaks, connection leaks, log/disk growth?Moderate load for 4-12+ hours
Scalability testDoes adding pods/CPU increase capacity proportionally?Same test on 1, 2, 4 nodes; compare max throughput
Volume testDoes it slow down with large data (10 crore rows)?Normal load against a production-sized database
Smoke perf testIs the script and environment sane before a big run?1-5 users, few minutes

The KPIs you report

  • Response time - time from sending the request to receiving the last byte. Report percentiles (p50/median, p90, p95, p99), not just the average. p95 = 2 s means 95% of requests finished in 2 s or less.
  • Latency / TTFB - time to the first byte; shows server think time separately from download time.
  • Throughput - completed requests or business transactions per second (TPS/RPS). Also network throughput in KB/s.
  • Error % - failed requests (HTTP errors, timeouts, failed assertions) divided by total.
  • Concurrency - active virtual users (VUs) or in-flight requests.
  • Resource utilisation - CPU, memory, GC, DB connections, thread pools on the servers.
  • Apdex - (satisfied + tolerating/2) / total, a 0-1 user-satisfaction score used in JMeter dashboards and APM tools.

Little's law: the formula interviewers love

For a stable system: N = X × (R + Z) where N = concurrent users, X = throughput (req/s), R = response time (s), Z = think time (s). Example: the business wants 100 searches/s, a search takes 0.5 s and users think for 4.5 s between actions. You need N = 100 × (0.5 + 4.5) = 500 virtual users. With zero think time, 50 users would produce the same 100 req/s - which is why removing think time makes a test unrealistically aggressive.

Writing a good NFR

“The site should be fast” is not testable. A good non-functional requirement is SMART: it names the transaction, the load, the percentile, the limit, the error budget, the duration and the environment. Ask the product owner or derive it from production analytics (peak hour from Google Analytics, access logs or APM).

NFR-PERF-01  Find Flights (POST /reserve.php)
  Load      : 200 concurrent users, 5 s think time, 30 min steady state
  Response  : p95 <= 2.0 s, p99 <= 4.0 s (measured at the load injector)
  Throughput: >= 35 transactions/s sustained
  Errors    : < 1 % (HTTP 4xx/5xx + failed assertions)
  Resources : app server CPU < 70 %, memory with no upward trend
  Data      : 10,000 bookings in DB before test; 50 unique routes in test data

Hands-on Lab: Measure percentiles, throughput and Little's law with real traffic

  1. Install Python 3.8+ (python3 --version). No libraries are needed.
  2. Break one request into phases with curl (Git Bash on Windows works too). Run it 3 times and note how DNS/TLS drop after the first call because of caching:
    curl -o /dev/null -s -w "dns=%{time_namelookup}s connect=%{time_connect}s tls=%{time_appconnect}s ttfb=%{time_starttransfer}s total=%{time_total}s code=%{http_code}\n" https://blazedemo.com/
    Expected: something like dns=0.01s connect=0.2s tls=0.45s ttfb=0.7s total=0.72s code=200. TTFB minus TLS is roughly server processing time.
  3. See why averages lie. Open a Python shell and run:
    >>> import statistics, math
    >>> times = [200] * 95 + [5000] * 5        # 100 requests, 5 very slow ones
    >>> statistics.mean(times)
    440
    >>> s = sorted(times)
    >>> s[math.ceil(0.95 * 100) - 1], s[math.ceil(0.99 * 100) - 1]
    (200, 5000)
    The average (440 ms) describes nobody: 95 users saw 200 ms and 5 users waited 5 s. p99 exposes them.
  4. Save the following as perf_basics.py and run python3 perf_basics.py:
    # perf_basics.py  -- run: python3 perf_basics.py   (Python 3.8+, no extra packages)
    import math
    import statistics
    import time
    import urllib.request
    from concurrent.futures import ThreadPoolExecutor
    
    URL = "https://blazedemo.com/"
    REQUESTS = 30      # keep it small: this is a shared public demo site
    WORKERS = 5        # 5 concurrent "users" with zero think time
    
    
    def hit(_):
        start = time.perf_counter()
        try:
            with urllib.request.urlopen(URL, timeout=15) as resp:
                resp.read()
                ok = resp.status == 200
        except Exception:
            ok = False
        return (time.perf_counter() - start) * 1000, ok
    
    
    def pct(values, p):
        """Nearest-rank percentile (same idea JMeter/k6 reports use)."""
        s = sorted(values)
        k = max(0, math.ceil(p / 100 * len(s)) - 1)
        return s[k]
    
    
    if __name__ == "__main__":
        t0 = time.perf_counter()
        with ThreadPoolExecutor(max_workers=WORKERS) as pool:
            results = list(pool.map(hit, range(REQUESTS)))
        wall = time.perf_counter() - t0
    
        times = [r[0] for r in results]
        errors = sum(1 for r in results if not r[1])
        avg_s = statistics.mean(times) / 1000
        throughput = len(times) / wall
    
        print(f"samples={len(times)} errors={errors} ({errors / len(times):.1%})")
        print(f"avg={statistics.mean(times):.0f} ms  median={pct(times, 50):.0f} ms")
        print(f"p90={pct(times, 90):.0f} ms  p95={pct(times, 95):.0f} ms  "
              f"p99={pct(times, 99):.0f} ms  max={max(times):.0f} ms")
        print(f"throughput={throughput:.2f} req/s over {wall:.1f} s")
        # Little's law: N = X * R  ->  predicted concurrency should be close to WORKERS
        print(f"Little's law check: X * R = {throughput * avg_s:.2f} (workers = {WORKERS})")
  5. Record the output in a table: avg, median, p90, p95, p99, max, error %, throughput. Expected order: median ≤ p90 ≤ p95 ≤ p99 ≤ max, and the Little's law line should print a value close to 5 (the number of workers), because with no think time N = X × R.
  6. Change WORKERS to 1 and then 10 (keep REQUESTS at 30). Observe: throughput rises with workers while the server is not saturated; if response time also rises sharply, you are seeing queueing.
  7. Using Little's law, calculate the VUs needed for 40 TPS on blazedemo if R = 0.8 s and Z = 4 s (answer: 40 × 4.8 = 192 VUs). Write the calculation in your notes.
  8. Write three NFRs for blazedemo (Home, Find Flights, Purchase) in the format shown above. You will reuse them in the capstone.

Ethics first: blazedemo.com and test.k6.io exist for learning, but they are shared. Keep public-site tests small (tens of users, minutes not hours). Heavy stress tests belong on your own environment - we will use a local Docker app in Lesson 6.

Common mistakes

  • Reporting only the average response time - always give p90/p95/p99 and the max.
  • Testing without agreed NFRs, then arguing about whether 3 s is “good”.
  • Removing think time to “generate more load” - it changes the user mix and inflates throughput unrealistically.
  • Mixing up response time measured at the client (includes network) with server time from APM.
  • Running a load test against production or a shared demo site without written permission.
  • Calling a 10-minute run a soak test - leaks need hours to show up.

Real-world assignment

Your manager forwards a mail: “Diwali sale expected 3x normal traffic. Normal peak = 1,200 orders/hour, average 6 page views per order, median page response 0.9 s, user think time 8 s. Tell us how many VUs to simulate and what NFRs we should sign off.” Produce a one-page note with: peak TPS (pages/s), required VUs via Little's law, a test-type plan (smoke, load, stress, spike, 4-hour soak) and five SMART NFRs. Hint: 3 × 1,200 × 6 / 3600 = 6 pages/s; N = 6 × (0.9 + 8) ≈ 54 VUs at the 3x peak, so plan the stress test to go well beyond that.

Key takeaways

  • Load = expected peak, stress = beyond peak, spike = sudden burst, soak = long duration, scalability = more hardware, more capacity?
  • Percentiles (p95/p99) describe user experience; averages hide outliers.
  • Little's law N = X × (R + Z) converts business throughput into virtual users.
  • Throughput, response time, error % and server utilisation must be read together.
  • Every test needs SMART NFRs and a target environment you are allowed to load.

🧪 Practical checklist

Do each task yourself and tick it off. All tasks are required to complete this lesson.

🎤 Interview questions

What is the difference between load, stress and soak testing?

A load test checks NFRs at expected peak load; a stress test pushes beyond peak to find the breaking point and failure behaviour; a soak test runs moderate load for hours to expose leaks and degradation over time.

Why do you report the 95th percentile instead of the average?

The average is distorted by outliers and hides the slow tail; p95 tells you the response time that 95% of users got or better, which maps directly to an SLA.

Explain Little's law with an example.

N = X × (R + Z). For 50 TPS with 1 s response time and 4 s think time you need 50 × 5 = 250 concurrent virtual users.

What makes a good performance NFR?

It names the transaction, load level, percentile, limit, error budget, duration and environment, e.g. 'Search p95 ≤ 2 s at 200 users with < 1% errors for 30 min'.

📝 Quiz, progress tracking & certificate

You're reading a free preview. Premium members tick off labs, take the quiz, unlock all 10 lessons and earn a verifiable certificate.

💎 Unlock with Premium Premium login
✓ You're subscribed! Job alerts arrive daily at 9 AM.
Scroll to Top