The First Thing I Validate Is the Load Generator
What NVIDIA's AIPerf rewrite clarified for me about the GIL, TTFT, and Poisson arrivals.
NVIDIA replaced GenAI-Perf with a multiprocess load generator. The change raises a basic benchmarking question: at high concurrency, are you measuring the inference server or the client generating the load?