Skip to learning content
Learning StudioCourses and practice
Experiences

Trace one inference request

Check prefill, decode, cache accounting, admission, and scheduling, then walk through the request timeline they produce.

This optional integration check reruns the module files together. Each lesson is complete only after its code, experiment, and knowledge check are complete.

Module files
systems/inference-runtime.pyIncomplete: code, experiment, checksystems/continuous-batching.pyIncomplete: code, experiment, check