Profiling¶
Profiling explains where a single run spends time or allocates memory. It is different from benchmarking: use profiling to find a candidate optimization and benchmarking to measure whether that optimization improved performance relative to a baseline.
PYDGENS provides optional profiling dependencies. Install them into the same Python environment used to run the example or tests:
Profile representative, focused workloads rather than the entire test suite. The commands below use the unicycle example; replace it with the example or test that exercises the code path of interest. Profilers add overhead, so do not use their elapsed times as benchmark results.
Scalene: CPU and Memory Hotspots¶
Scalene is useful for a broad line-level view of CPU and memory activity. The command below profiles the unicycle example and saves its data outside the source tree:
mkdir -p .profiles
scalene run --profile-all \
--outfile .profiles/unicycle-scalene.json \
src/pydgens/examples/unicycle.py
Open the interactive report or render a terminal summary:
--profile-all includes code beyond the target script. For a faster CPU-only
pass, replace it with --cpu-only. To profile a focused pytest workload,
Scalene can run pytest as a module:
Line Profiler: One Function at a Time¶
Line Profiler measures execution time line by line for functions you explicitly mark. It has higher overhead than a sampling profiler, so use it after another tool has identified a small area to investigate.
Temporarily decorate the function of interest. Importing the decorator keeps the module runnable when it is not being profiled:
Run the module through kernprof with line-by-line profiling enabled. The
-v flag prints the report and -o saves it for later inspection:
mkdir -p .profiles
PYTHONPATH=src kernprof -l -v \
-o .profiles/unicycle.lprof \
-m pydgens.examples.unicycle
View a saved report again with:
Remove temporary decorators before committing unless they are intentionally part of the maintained profiling setup.
PyInstrument: Call-Tree Overview¶
PyInstrument is a sampling profiler that produces a call-tree overview with relatively low overhead. It can profile a script or module directly, without modifying the source code:
Save an interactive HTML report for sharing or later inspection:
mkdir -p .profiles
pyinstrument \
--renderer html \
--outfile .profiles/unicycle-pyinstrument.html \
-m pydgens.examples.unicycle
To profile selected tests, run pytest through PyInstrument:
For a very short workload, PyInstrument may collect too few samples. In that case, profile a representative loop or use a smaller sampling interval, while recognizing that smaller intervals add overhead.
Interpreting Results¶
Treat profiler output as evidence for where to investigate, not as a direct performance claim. JAX compilation and asynchronous execution can make a first-run profile look very different from a steady-state solve. After making an optimization, use the benchmark comparison procedure to compare the same workload against a baseline under controlled conditions.