nvbench

mirror of https://github.com/NVIDIA/nvbench.git synced 2026-06-29 10:47:36 +00:00

Author	SHA1	Message	Date
Oleksandr Pavlyk	f18d48d2fd	Ensure that bulk-debug-python script is enclosed in markers This permits extracting Python script using Unix CLI tools when `--bulk-debug-python stdout` is used. Added example of using this to nvbench_compare.md doc.	2026-06-22 11:58:16 -05:00
Oleksandr Pavlyk	5cc14fc948	Replaced UNDECIDED with AMBG, use Gray color/shrug emoji	2026-06-22 09:55:42 -05:00
Oleksandr Pavlyk	498ef45247	Improve nvbench-compare interval display readability Add compact reason labels for explain-mode tables while keeping canonical reason codes in the undecided summary. Emit a one-line legend only for non-trivial abbreviations. Refine interval displays so timing values align across table rows: - align Lo/Ce/Hi values in explain mode - align center values in intervals mode when some rows lack interval bounds - avoid repeating units when center and interval deltas use the same unit Add a Change column for non-legacy displays so FAST/SLOW rows show the signed interval-bound relative change, while SAME and UNDECIDED rows remain blank. Extend nvbench_compare tests to cover reason legend filtering, interval alignment, missing-interval alignment, and Change column formatting.	2026-06-04 15:33:13 -05:00
Oleksandr Pavlyk	9890aad294	Implement --bulk-debug-python option Use this option to generate Python script with information needed to load bulk data from reference/compare datasets for further drill-down into data.	2026-06-04 12:49:55 -05:00
Oleksandr Pavlyk	997e0be9db	Rename IR to IQR in /ir/absolute and /ir/relative tags	2026-06-04 11:16:53 -05:00
Oleksandr Pavlyk	9f67f15684	Add document for nvbench_compare It documents use, documents decision tree, configurability, use of --display option, benchmark/axis filtering and device filtering	2026-06-04 11:15:18 -05:00
Oleksandr Pavlyk	a385ee5335	Support rename of tags /ir/(absolute\|relative) to /iqr/(absolute\|relative)	2026-06-04 11:15:10 -05:00
Oleksandr Pavlyk	841bd87638	Add TOML configuration for nvbench-compare thresholds Add versioned TOML configuration support for nvbench-compare threshold settings. The new --config option reads grouped settings for clear-gap, same-result, bulk coverage, and rare-support filtering thresholds. The parser validates the schema strictly so unknown tables, unknown keys, invalid types, unsupported versions, and out-of-range values fail early. Add --dump-config to print the effective configuration without requiring input JSON files. This makes the currently selected preset and resolved threshold values discoverable and gives users a starting point for custom configuration. Preset resolution is: - default is used when neither TOML nor CLI selects a preset - [preset] name = "..." in TOML selects the base preset - --preset ... overrides the TOML preset selection - explicit threshold values in TOML override whichever base preset was selected For example: - nvbench-compare --dump-config Prints the built-in default settings as grouped TOML. - nvbench-compare --preset permissive --dump-config Prints the permissive preset values as TOML. - nvbench-compare --config compare.toml ref.json cmp.json Compares using the preset named in compare.toml, plus any explicit TOML threshold overrides. - nvbench-compare --config compare.toml --preset strict ref.json cmp.json Uses the strict preset as the base, while preserving explicit threshold overrides from compare.toml. Keep TOML parsing lazy: Python 3.11+ uses tomllib, while Python 3.10 only requires tomli when --config is used. Add focused tests for grouped config dumping, strict validation, preset/override precedence, and CLI dump behavior.	2026-06-04 09:55:58 -05:00
Oleksandr Pavlyk	4cf75dcaf5	Add nvbench_compare display modes and interval-based table views Extend nvbench_compare with multiple table display modes and richer interval formatting for timing comparisons. Highlights: - add `--display` with `intervals`, `legacy`, and `explain` modes - keep `legacy` output using scalar Diff/%Diff - make `intervals` the default, showing compact center-plus-delta timing intervals - add `explain` mode with explicit `[L \| C \| H]` interval rendering and self-describing headers - compute and store diff and relative-diff intervals in SummaryComparison - add formatting helpers for absolute and relative interval displays - make default preset slightly more permissive by lowering `bulk_same_sample_coverage` to 0.97 Add focused tests covering: - diff/%diff interval computation - compact and explicit interval formatting - default, legacy, and explain table layouts - CLI propagation of `--display` and preset selection	2026-06-04 08:49:06 -05:00
Oleksandr Pavlyk	2a515c2569	Change in how FAST/SLOW deciision is arrive at Now: - establish a candidate clear timing gap from summary timing intervals, as before - if bulk sample times and frequencies are available on both sides, compute cycles = time * frequency - derive bulk cycle intervals from min/q1/median/q3 - confirm the gap direction from those bulk cycle intervals - only fall back to summary sm_clock_rate_mean confirmation when bulk cycle data is unavailable I also split the reason codes so the evidence source is visible: - clear_gap_confirmed_by_bulk_cycles - bulk_cycle_gap_not_confirmed - clear_gap_confirmed_by_summary_cycles - summary_cycle_gap_not_confirmed Updated tests in python/test/test_nvbench_compare.py cover both the bulk-confirmed and bulk-rejected paths, along with the renamed summary reason codes.	2026-06-03 15:57:34 -05:00
Oleksandr Pavlyk	20b3bd3148	Add nvbench_compare presets and rare-support-aware bulk coverage Introduce comparison threshold presets in nvbench_compare and thread the selected preset through main() into compare_benches. Refine bulk nearest-neighbor support handling by: - adding rare-support filtering thresholds - ignoring low-count support values only when removed sample mass is small - falling back to full support for all-unique or otherwise unusable support - keeping sample-weight coverage over all values Tighten bulk mismatch reporting to show compact min(ref, cmp) coverage summaries, and add tests covering: - rare-tail filtering - strict fallback when too much support mass would be removed - all-unique support preservation - preset lookup and CLI preset propagation	2026-06-03 15:21:26 -05:00
Oleksandr Pavlyk	b791522d48	Group nvbench-compare thresholds into a config object Replace the scattered module-level comparison threshold constants with a ComparisonThresholds value object. Thread this object through compare_benches, compare_gpu_timings, and the lower-level clear-gap, summary-SAME, and bulk-SAME decision helpers. Keep existing behavior by constructing default ComparisonThresholds when callers do not provide one. This prepares nvbench-compare for future CLI-configurable decision thresholds while keeping one consistent configuration for an entire comparison run. Add test coverage that passes custom thresholds through compare_benches and verifies they affect the SAME decision.	2026-06-03 10:02:46 -05:00
Oleksandr Pavlyk	8c85393ee2	Use bulk samples to confirm same comparisons Add a bulk-data SAME path to nvbench_compare for cases where summary intervals do not provide a clear FAST/SLOW decision. The new path compares sample times and SM-clock-adjusted cycles with symmetric nearest-neighbor coverage over unique values and sample counts. The comparison now requires both sample-weight coverage and unique-support coverage to pass before declaring SAME. If bulk data is available but coverage does not pass, the result remains UNDECIDED instead of falling back to the summary-only SAME rule. Also improve undecided diagnostics by aggregating reason codes while preserving the most severe representative detail, including observed coverage values and thresholds for bulk support mismatches. Add tests for: - bulk data confirming SAME despite changed mode weights; - bulk time mismatch overriding summary-only SAME; - cycle coverage vetoing time-only agreement; - sample-weight and unique-support coverage diagnostics; - aggregation of undecided reason details.	2026-06-03 09:36:05 -05:00
Oleksandr Pavlyk	65abfbcfb2	Implement DecisionReason, tracking and summarisation - Add DecisionReason(code, message) and internal TimingDecision(status, reason). - SummaryComparison now carries reason - ComparisonStats now aggregates undecided reasons. - Final summary prints a reason breakdown only when undecided reasons exist, e.g.: - Undecided (comparison requires more evidence): 3 - Reasons: - noise_too_high: 2 (relative dispersion is too high to declare same) - weak_interval_overlap: 1 (timing intervals do not overlap strongly enough to declare same)	2026-06-03 07:52:25 -05:00
Oleksandr Pavlyk	6de54fa07a	Implement early SAME check If SLOW/FAST check returned undecided, we attempt conservative SAME check based on summary data alone (bulk data are not read) Reference and compare measurements are considered SAME if - both centers are positive finite values; - abs(ref - cmp) / min(ref, cmp) <= 0.5%. This is equivalent to max(ref, cmp) / min(ref, cmp) <= 1 + delta; - interval overlap must cover at least 50% of the smaller interval; - relative dispersion must be finite on both sides and no more than 2%; - if SM clock summaries are available, the same check must also pass in cycle space. Otherwise UNDECIDED remains working decision, to be refined by further checks	2026-06-03 07:38:00 -05:00
Oleksandr Pavlyk	48b7f61da3	Implement clear-gap comparison for early FAST/SLOW decision Implemented the clear-gap comparison, with the log-distance-equivalent algebra and pessimistic SM-clock fallback. What changed: - Added TimingInterval and interval construction from summaries: - robust interval: [min, q3], centered at median - fallback interval: clipped [mean - stdev, mean + stdev] intersected with [min, max] - Added CLEAR_GAP_RELATIVE_THRESHOLD = 0.005. - FAST gap uses: (ref.lower - cmp.upper) / cmp.upper >= delta which is equivalent to log(ref.lower / cmp.upper) >= log(1 + delta). - SLOW gap uses: (cmp.lower - ref.upper) / ref.upper >= delta - FAST/SLOW now requires SM clock summaries on both sides and the same clear-gap result after scaling intervals by sm_clock_rate_mean. - If intervals are missing, overlap, fail the gap threshold, have missing/invalid clock summaries, or time/cycle comparison disagrees, status is UNDECIDED. - Existing center/noise values are still computed and displayed, but no longer drive FAST/SLOW/SAME classification. Updated tests to cover: - center/noise-only comparisons becoming UNDECIDED - clear FAST/SLOW with matching clock evidence - missing clock fallback to UNDECIDED - frequency-shift disagreement becoming UNDECIDED - regression reporting with robust interval and clock evidence	2026-06-03 07:13:46 -05:00
Oleksandr Pavlyk	71823e2f4f	Add q1/q3 quartiles to GPUTimeData struct The quantile values are not currently used, but plumbed through	2026-06-03 06:35:24 -05:00
Oleksandr Pavlyk	0d1d9d2838	Merge remote-tracking branch 'upstream/main' into nvbench-compare-process-bulk-data	2026-06-02 17:02:57 -05:00
Oleksandr Pavlyk	a8704103a7	Add "nv/cold/sm_clock_rate/mean" to GPU time summary data Its intent is to be cheaply retrievable metric of average SM clock frequence over entire sample	2026-06-02 16:21:39 -05:00
Oleksandr Pavlyk	debde4f4b2	Lazy-load nvbench-compare bulk timing data Store JSON-bin sample time and frequency metadata in GpuTimingData instead of reading the binary files during summary extraction. Add Float32BinarySource and lazy cached accessors for samples and frequencies. Use np.fromfile by default, but allow tests and alternate callers to inject a float32 reader returning any buffer-compatible object convertable to "<f4" data type. Treat optional bulk-data failures as unavailable evidence instead of aborting comparison: unreadable files, invalid buffers, count mismatches, and mismatched sample/frequency metadata now emit RuntimeWarning and return None. Update nvbench_compare tests to verify lazy loading, cache reuse, injected reader behavior, warning-based degradation, and count mismatch handling.	2026-06-02 15:55:02 -05:00
Oleksandr Pavlyk	6d8aa878cf	Introduce UNDECIDED comparison status It is not emitted just yet, but the code becomes ready for it when it starts being emitted	2026-06-02 15:23:47 -05:00
Oleksandr Pavlyk	d4283f77a5	Refactor nvbench-compare timing comparison state Introduce GpuTimingData, SummaryComparison, ComparisonStats, and ComparisonRunData to make timing extraction, classification, and run-level state explicit. Load sample-time and SM-frequency bulk data from JSON binary output into GpuTimingData when available, preserving count validation between paired sample and frequency arrays. Move GPU timing comparison logic into compare_gpu_timings(), prefer robust median/IQR data when available, and fall back to mean/stdev summaries otherwise. Keep missing or invalid noise on the unknown path. Replace module-level comparison counters and selected-device globals with per-run data passed into compare_benches(). Update tests to validate timing classification, bulk-data loading, device pairing, filtered duplicate matching, and summary counters through the new structures.	2026-06-02 15:04:39 -05:00
Oleksandr Pavlyk	0b2dd26625	Make nvbench_compare read bulk data, if available	2026-06-02 13:38:53 -05:00
Oleksandr Pavlyk	0dc93b0c0e	Introduce robust metrics (#379 ) * Add statistics utilities to compute quartiles Quartiles are computed using nearest rank method. Two implementations are provided: 1. Sort-based: a. sort array b. extract values at ranks of interest 2. Selection based: a. Run nth_element to find median on whole range b. Run nth_element on left side to find first quartile c. Run nth_element on right side to find thirst quartile Public API copies input into temporary vector which is mutated as needed. Public API uses sort-based implementation for small arrays ( <= 4096 elements), and selection-based implementation for larger arrays. Sort-based implementation can support computation of arbitrary percentiles, which could be useful later if more extreme statistics is needed. Add tests covering percentile and quartile edge cases, input iterators, selection-vs-sorting agreement, empty and singleton inputs, and relative dispersion validation. * Add quartiles information to summaries Use the quartile helpers to report robust cold and CPU-only timing summaries: Q1, median, Q3, interquartile range, and relative interquartile range. These values stay hidden. Summary tags are nv/cold/time/gpu/q1, nv/cold/time/gpu/median, nv/cold/time/gpu/q3, nv/cold/time/gpu/ir/absolute, nv/cold/time/gpu/ir/relative ir/absolute = q3 - q1, ir/relative = (q3 - q1)/median Similar tags added for nv/cold/time/cpu and for CPU-only measures. Validate relative-dispersion calculations before publishing relative noise summaries so invalid centers or dispersion values do not produce misleading summary entries. * Prefer robust summaries in default output Only flip visibility for nv/cold/cpu/time, nv/cold/gpu/time, and nv/cpu_only/only: - hide mean - hide stdev/relative - show median - show ir/relative * Use is_close where std::abs(act-exp) was used * Revert "Prefer robust summaries in default output" This reverts commit `9a0afc361c`. Basically, all robust statistics summaries entries are hidden, and mean + stdev/relative are back to be default displayed items * Address PR review feedback	2026-06-02 13:20:15 -05:00
Oleksandr Pavlyk	1d13b49996	Add scoped filtering and device pairing to nvbench_compare Teach nvbench_compare to keep the order of --benchmark and --axis arguments so axis filters can apply either globally or to the most recent benchmark. Build a filter plan from the ordered CLI arguments and apply the same plan to table output and plotting labels. Add explicit --reference-devices and --compare-devices filters. The filters accept all, a single device id, or a comma-separated list of ids; ordered lists and duplicates are preserved so selected reference and compare devices can be paired by position. Device-section mismatches remain fatal for unfiltered all-vs-all comparisons, but become warnings when the user explicitly selects devices and the selected device counts match. Match duplicate benchmark states by occurrence within each filtered device section instead of matching only by state name across the whole benchmark. This keeps repeated axis values and filtered duplicate states aligned between the reference and compare inputs, and reports mismatched occurrence counts instead of silently dropping extra states. Add Python tests for duplicate-state matching, axis filtering before matching, device filter parsing and validation, explicit cross-device pairing, and benchmark-scoped axis filters. Original commit messages folded into this change: Tweaks for nvbench_compare 1. When JSON files contain multiple entries with the same name and axis values, make sure that scripts compares corresponding entries. Previous logic would extract the first entry from ref data, and would compare measurements for each state in cmp against the first entry from ref. The change introduces a counter to know which nth entry we process for a particular axis value, and retrieve corresponding entry in ref. Scope occurrence matching by device. Device pairing in nvbench_compare.py is strictly index-based under --ignore-devices, reused IDs in a different order no longer pair against the wrong reference device. Require devices in ref and cmp to have the same cardinality Handle mismatch when number of duplicates in ref data is not same as in cmp data Use pytest monkeypatch fixture to pretend third-party package dependencies are available during test run for nvbench_compare without introducing test-time dependency Added the happy-path test and fixed its direct-call setup by initializing the device globals that main() normally populates. Fix to filter-before-matching. - compare_benches() now pairs devices by selected position instead of taking a device id. - For each device pair, compare_benches() now builds: - ref_device_states: matching reference device and axis filters - cmp_device_states: matching compare device and axis filters - State occurrence counts and duplicate occurrence matching now operate only on those filtered per-device lists. - Removed the later matches_axis_filters() skip inside the compare-state loop because filtering now happens before matching. Added a regression test where ref/cmp have duplicate state names in opposite order, and --axis keeps only one of them. The test verifies the kept compare state is matched against the kept reference state, not the first unfiltered occurrence. Introduce device filtering in nvbench_compare - --reference-devices all\|ID\|ID,ID,... - --compare-devices all\|ID\|ID,ID,... - Integer lists preserve order and duplicates. - Requested IDs are validated against the file-level device list. - Filtered reference/compare device counts must match before comparison. - compare_benches() pairs selected reference and compare devices by position. - Each benchmark validates that requested device IDs are present in its own devices list. Implemented benchmark-scoped --axis handling. - --axis and --benchmark now share an ordered argparse action, so their relative CLI order is preserved. - -a before any -b becomes a global axis filter. - -a after -b <name> applies to that most recent benchmark only. - Repeated -b entries are treated as separate filter scopes and combined as alternatives for that benchmark. - Device filtering remains global and is applied independently. Allow non-matching devices for explicit device selection Now the device-section equality check remains fatal only for unfiltered all-vs-all comparisons. If either --reference-devices or --compare-devices is explicit, mismatched selected device metadata is printed as a warning, but comparison proceeds after the selected device counts have been validated. Fix for resolve_benchmark_device_ids, add comments The return value of resolve_benchmark_device_ids now always owns its list. Use monkeypatch class in set_test_devices helper Stricted device id validation Test for device id validation	2026-06-02 11:48:01 -05:00
Oleksandr Pavlyk	ca1d60610c	Use robust summaries in nvbench_compare classification Teach nvbench_compare to parse GPU timing summaries into structured values and prefer the robust median/IQR summaries when both compared measurements provide them. Fall back to the existing mean/stdev summaries when robust summaries are not available. Classify comparisons with the larger available relative noise estimate instead of the smaller one, keep unavailable noise distinct from encoded infinite noise, and report improvements separately from regressions. Keep the process exit code as success for completed comparisons; regression counts are reported in the summary instead of being used as the process status. Make plotting tolerate unavailable noise by leaving gaps in confidence bands, sort plotted series by the plotted axis, and avoid reusing pyplot state across plot calls. Add focused Python tests for robust-summary preference, unavailable-noise classification, non-finite timing centers, plot-along handling when the selected axis is absent, and the exit-code contract.	2026-06-02 11:47:47 -05:00
Oleksandr Pavlyk	ee4b9f0963	Remove unused python_wheel section (#382 ) ci/matrix.yaml contains unused section once intended for Python wheels	2026-06-01 14:04:38 -05:00
Oleksandr Pavlyk	97c8b29f5a	Updated devcontainer imageset to 26.08 (#381 ) Add CTK 13.2 with compact support for host compilers: - gcc 11 (min), gcc 13 (working), gcc 15 (max) - llvm15 (min), llvm 21 (max) - CL 14.44	2026-06-01 11:02:40 -05:00
Oleksandr Pavlyk	7ba2b79d5b	Reduce stdrel criterion complexity and ensure termination (#374 ) * Reduce stdrel criterion complexity and ensure termination Replace the stdrel criterion's growing sample history with an online mean/variance accumulator. This keeps the stopping criterion based on relative standard deviation, preserves the unbiased standard-deviation estimate used for convergence, and reduces per-sample update work from recomputing over the full history to constant time. Add a bounded invalid-noise path so measurements that persistently produce non-finite relative noise, such as all-zero timings, can terminate without waiting for the wall-time timeout. Keep the normal min-time gate for ordinary stdrel convergence. Add focused tests for the online accumulator, stdrel sample-count threshold, sample-standard-deviation behavior, deterministic convergence inputs, and persistent invalid-noise termination. Update the CLI help for the stdrel termination behavior. * change max-noise to for consistency * Use online_mean_variance on m_noise_tracker in is_finished() Previously, standard deviation call was made using current noise level instead of mean noise level. Because of identity E[ (N - C)^2 ] = E[ (N - E[N])^2 ] + (E[N] - C)^2 >= E[ (N - E[N])^2 ] this led to criterion terminating later than it could have because the estimated expectation is always greater of equal that the estimate relative to the mean. Code used current noise level instead of mean to avoid needing to make two passed through m_noise_tracker container. Use of online_mean_variance allows to improve accuracy of estimating dispersion of noise signal while maintaining single pass through container. * Address review feedback Fixed misleading commit. Introduce private methods to refactor computation of repeated expressions. Renamed m_cuda_times_summary to m_measurements_summary, since criterion can be applied for CPU-only measurements too. Introduced is_close utility for checking whether two floating point numbers are closed to one another. Introduced descriptive constexpr variables for hard-wired constants	2026-05-29 17:06:28 +00:00
omribz156	ec025d7e0d	docs: separate measurement options from stopping criteria (#373 ) Signed-off-by: Omri SirComp <omribz156@gmail.com>	2026-05-28 16:51:12 -05:00
Oleksandr Pavlyk	6bdbff7f21	include cleanup across nvbench/ (#377 ) Added missing direct standard includes for entities such as std::size_t, std::move, std::vector, std::optional, std::exception, std::memcpy, etc. Added missing project include in nvbench/internal/table_builder.cuh for nvbench::detail::transform_reduce. Fixed nvbench/detail/gpu_frequency.cuh to forward-declare nvbench::cuda_stream in nvbench namespace instead of in nvbench::detail namespace.	2026-05-28 16:40:30 -05:00
Oleksandr Pavlyk	84c7952f8b	nvbench::cpu_timer changed to use steady_clock (#371 ) Using steady_clock is more appropriate for timing measurements. It guarantees that duration computed from two time-points will not contain correction deltas.	2026-05-20 10:22:22 -05:00
mfranzrebsal	4a33a61591	Add Windows support (#354 )	2026-05-19 15:10:58 -05:00
Oleksandr Pavlyk	3d82e58170	Fix docutil error when building docs (#365 )	2026-05-18 10:57:19 -05:00
Oleksandr Pavlyk	4472e7b59b	Add python api for cold warmup parameters (#363 )	2026-05-18 10:56:44 -05:00
Oleksandr Pavlyk	ce75dab94b	Add stopping criterion sample count (#341 ) * Implement sample-count stopping criterion with parameter target-samples --stopping-criterion sample-count --target-samples 100 would stop once max(--min-samples, --target-samples) samples are collected * Address review nitpicks	2026-05-15 15:15:12 -05:00
Oleksandr Pavlyk	6dd27aedfd	Fix exception safety (#358 ) Improve exception safety of timer structs by using local scope guards to ensure that cleanup steps, such as signaling blocking kernel to unblock and making sure that the stream is synchronized are performed even launch object throws an exception. Tests of exception safety were added. -- * blocking_kernel.unblock_noexcept() noexcept method added This decouples the logic of signaling to unblock from checking of the timeout. * Improve exception safely in kernel_launch_timer Introduce noexcept cleanup methods. Place body of start() and stop() methods in the try/catch block and execute noexcept clean-up on exception before rethrowing. * Improve exception safety of measure_hot * Make sure that throwing methods call noexcept ones instead of duplicating functionality * Use cleanup_guard in measure_cold_base::kernel_launch_timer Replace try/catch pattern with cleaner use of cleanup_guard class. * cpu_timer::start, cpu_timer::stop methods marked noexcept These methods do not throw, and marking them noexcept explicitly makes it fine to call them from other noexcept methods, as such cleanup_noexcept in measure_cold. * Address remaining exception safety issue in measure_hot * Renamed guard variables to reflect their purpose, apply arm-then-do to ops queueing kernels Set m_block_stream_armed = true; before launching the kernel. Doing so signals cleanup guard that stream must be unblocked, even if launching of the kernel failed. Same for operation launching time-stamps kernel. * Add testing/device/exception_safety.cu This test add benchmark that throws. It verifies that it did not time-out and control counters the benchmark maintains are at the expected values. * Refactor measurement cleanup guards for testability Extract hot stream cleanup and cold launch timer cleanup into reusable detail helpers. Keep measure_hot and measure_cold using those helpers through thin adapters so the tested cleanup logic matches the production path. Add driver-free cleanup guard tests using a fake measure object to verify cleanup ordering when exceptions occur after blocking stream setup, after hot unblock, and around cold GPU frequency start/stop paths. * Implement cpu_timer_stop_noexcept in terms of cpu_timer_stop The cpu_timer_stop is already noexcept by nature of implementation, but we maintain cpu_timer_stop_noexcept method for symmetry with other pairs sync_stream()/sync_stream_noexcept(). The cpu_timer_stop_noexcept() is implemented via cpu_timer_stop(). These methods are annotated __forceinline__, so the same code should be generated. * More readable initialization of bool members * Moved exception_safety.cu back to testing/ folder testing/device is reserved for tests that require locking of GPU frequency per CMake option description. * Fixed nitpick and bug it discovered Changed testing/exception_safety.cu:237 so run_benchmark now iterates over every state from bench.get_states() and checks each one is skipped with a reason containing "requested". That exposed a real runner behavior gap, so I also made a minimal fix in nvbench/runner.cuh:120: after stop_runner_loop, remaining states are now explicitly marked skipped with a reason instead of only printing a skip notification. * Move static assertions (pertaining to cleanup guards) to testing/cleanup_guards.cu The CI failure with CTK 12.0 and certain version of GCC is caused by OOM in cudafe++ process tripped by compiling instantiation of contract verification on cold_launch_timer_probe struct. As a work-around, this instantiation is excluded for CTK 12.0-12.6	2026-05-15 15:14:30 -05:00
Oleksandr Pavlyk	d63a2761eb	Implement Timer, and support State.exec(fn, timer=True) (#364 ) * Add type annotations for future functionality ```python class Timer: def start(self) -> None: ... def stop(self) -> None: ... ``` and overloaded `State.exec` so: - normal mode accepts `Callable[[Launch], None]` - `timer=True` accepts `Callable[[Launch, Timer], None]` No implementation yet. Type annotation checked with ``` (py313) :~/repos/nvbench/python$ python -m mypy --ignore-missing-imports /tmp/check_timer.py /tmp/check_timer.py:24: error: No overload variant of "exec" of "State" matches argument types "Callable[[Launch], None]", "bool" [call-overload] /tmp/check_timer.py:24: note: Possible overload variants: /tmp/check_timer.py:24: note: def exec(self, Callable[[Launch], None], /, , batched: bool \| None = ..., sync: bool \| None = ..., timer: Literal[False] = ...) -> None /tmp/check_timer.py:24: note: def exec(self, Callable[[Launch, Timer], None], /, , timer: Literal[True], sync: bool \| None = ...) -> None /tmp/check_timer.py:25: error: Argument 1 to "exec" of "State" has incompatible type "Callable[[Launch, Timer], None]"; expected "Callable[[Launch], None]" [arg-type] /tmp/check_timer.py:26: error: No overload variant of "exec" of "State" matches argument types "Callable[[Launch, int], None]", "bool" [call-overload] /tmp/check_timer.py:26: note: Possible overload variants: /tmp/check_timer.py:26: note: def exec(self, Callable[[Launch], None], /, , batched: bool \| None = ..., sync: bool \| None = ..., timer: Literal[False] = ...) -> None /tmp/check_timer.py:26: note: def exec(self, Callable[[Launch, Timer], None], /, , timer: Literal[True], sync: bool \| None = ...) -> None Found 3 errors in 1 file (checked 1 source file) (py313) :~/repos/nvbench/python$ nl -ba /tmp/check_timer.py 1 # /tmp/check_nvbench_timer.py 2 import cuda.bench as bench 3 4 def normal_ok(launch: bench.Launch) -> None: 5 pass 6 7 def timer_ok(launch: bench.Launch, timer: bench.Timer) -> None: 8 timer.start() 9 timer.stop() 10 11 def missing_timer(launch: bench.Launch) -> None: 12 pass 13 14 def extra_timer(launch: bench.Launch, timer: bench.Timer) -> None: 15 pass 16 17 def wrong_timer_type(launch: bench.Launch, timer: int) -> None: 18 pass 19 20 def state_bench(state: bench.State) -> None: 21 state.exec(normal_ok) 22 state.exec(normal_ok, timer=False) 23 state.exec(timer_ok, timer=True) 24 state.exec(missing_timer, timer=True) # should fail 25 state.exec(extra_timer) # should fail 26 state.exec(wrong_timer_type, timer=True) # should fail ``` * Implement cuda.bench.Timer object The Timer class is not user-constructible. It exposes two nullary methods timer.start() and timer.stop(). The instance of Timer class would be provided to launchable object passed to State.exec with timer=True. * Implement support for State.exec( launch_fn, timer=True) * Change type annotation for batch to default to None None is interpreted as `not timer`, i.e., it effectively defaults to True (as before) for usage without timer set, but starts defaulting to `False` is `timer=True` is set. The batched keyword type is `bool \| None`. * Implement default batched=None behavior API allows one to specify all 3 keywords, sync, batched, and timer. batched is None by default, run-time interpreted as `(not timer)`. * Update tests for new behavior of batched/time combination * Add python/examples/exec_tag_timer.py * Expand Timer class and methods docstrings * Reworked python/example/exec_tag_timer.py to align with C++ example. * Replace ::cuda::std::name with cuda::std::name * Resolve review feedback	2026-05-15 10:19:40 -05:00
Oleksandr Pavlyk	44ec7de6bd	Implement decorators to register benchmarks add axis and options (#347 ) * Add decorators for registering benchmarks and adding axis cuda.bench.register(fn) continues returning Benchmark, and supports legacy use. New signature added: cuda.bench.register(): Returns a decorator ``` @bench.register() @bench.axis.float64("Duration (s)", [7e-5, 1e-4, 5e-4]) @bench.option.min_samples(120) def single_float64_axis(state: bench.State): ... ``` * Remove example/auto_throughput.py The C++ counterpart's purpose is to demonstrate use of CUPTI metrics, but these are not supported in Python bindings, so this example is a duplicate of example/throughput.py * Add wrong decorator order test for bench.axis.* * Strengthen type annotation for register function Acting on code rabbit nit-pick require that function being registered take cuda.bench.State object as an argument. Verified the fix as ``` (py313) :~/repos/nvbench/python$ python -m mypy --ignore-missing-import /tmp/t.py /tmp/t.py:8: error: Argument 1 has incompatible type "Callable[[], None]"; expected "Callable[[State], None]" [arg-type] Found 1 error in 1 file (checked 1 source file) (py313) :~/repos/nvbench/python$ nl -ba /tmp/t.py 1 # /tmp/check_nvbench_register.py 2 import cuda.bench as bench 3 4 @bench.register() 5 def good(state: bench.State) -> None: 6 pass 7 8 @bench.register() 9 def bad() -> None: 10 pass ``` * Replace use of global variable with thread-safe lru_cache This improves thread-safety of module initialization. * Abide by RUF005 linting rule * Expand docstrings regarding cuda.bench.register() decorator It explains to the user what the decorator does and provides a concise usage example. * Sharpen wording on exception maybe-thrown by decorator	2026-05-14 15:41:30 -05:00
Oleksandr Pavlyk	338936b6fe	Provide BenchmarkResult class for parsing JSON output of NVBench-instrumented benchmarks (#356 ) Implements `cuda.bench.results.BenchmarkResult` class to represent data from JSON output of benchmark execution. The contains implements two class methods `BenchmarkResult.from_json(filename : str \| os.PathLike, , metadata : Any = None)` which expects well-formed JSON filename and `BenchmarkResult.empty(, metadata : Any = None)` intended to represent failed result with reasons that can be recorded in metadata at user's discretion. The `BenchmarkResult` implements mapping interface, supporting `.keys()`, `.values()`, `.items()` methods, `__len__`, `__contains__`, `__getitem__` and `__iter__` special methods. Values in `BenchmarkResult` has type `cuda.bench.results.SubBenchmarkResult` which implements a list-like interface, i.e. implements `__len__`, `__getitem__`, and `__iter__` special methods. Values in this list-like structure correspond to measurements of individual states of a particular benchmark (the key in `BenchmarkResult`). Elements of `SubBenchmarkResult` structure have type `SubBenchmarkState` that supports mapping protocol with axis_values as a key and represent data corresponding to measurements for a particular state (combination of settings for each axis). The state provides `.samples` and `.frequencies` attributes storing raw execution duration values and estimates for average GPU frequencies. Example usage: ``` import array, numpy as np, cuda.bench.results r = cuda.bench.results.BenchmarkResult("perf_data/axes_run1.json") r["copy_sweep_grid_shape"].centers_with_frequencies( lambda t, f: np.median(np.asarray(t)np.asarray(f))) ``` ``` In [1]: import array, numpy as np, cuda.bench.results In [2]: r = cuda.bench.results.BenchmarkResult("temp_data/axes_run1.json") In [3]: list(r) Out[3]: ['simple', 'single_float64_axis', 'copy_sweep_grid_shape', 'copy_type_sweep', 'copy_type_conversion_sweep', 'copy_type_and_block_size_sweep'] In [4]: r["simple"].centers(lambda t: np.percentile(t, [25,75])) Out[4]: {'Device=0': array([0.00100966, 0.00101299])} In [5]: r.centers(lambda t: np.percentile(t, [25,75]))["simple"] Out[5]: {'Device=0': array([0.00100966, 0.00101299])} In [6]: len(r) Out[6]: 6 In [7]: "fake" in r Out[7]: False ``` Each `SubBenchmarkState` implements `.summaries` attribute - rich object that retains tag/name/hint/hide/description metadata. Add nvbench-json-summary to render NVBench JSON output as an NVBench-style markdown summary table, including axis formatting, device sections, hidden summary filtering, and summary hint formatting. Update packaging, type stubs, and tests for the new namespace, renamed classes, Python 3.10-compatible annotations, and summary-table generation. * Split tests in test_benchmark_result into smaller tests * Fix break due to file name change * Add python/examples/benchmark_result_autotune.py This example demonstrates using cuda.bench and cuda.bench.results to implement simple auto-tuning, demonstrated on selecting of tile shape hyperparameter for naive stencil kernel implemented in numba-cuda. * Resolve ruff PLE0604 * Fix for format_axis_value in json format script to handle None value Add tests to cover such input. * Address code rabbit review feedback * Fix license header, add validation * Addressed both issues raised in review Malformed values are now represented in result as None. Skipped benchmarks are no longer dropped, i.e., they are present in BenchmarkResult data, but they are not reflected in summary table in line with what NVBench-instrumented benchmarks do.	2026-05-13 13:23:58 -05:00
Oleksandr Pavlyk	6df6dc8d89	Enable building of NVBench on Windows (#362 ) * Enable building of NVBench on Windows, no testing * Add guard to disable nvbench-windows for now	2026-05-13 13:16:41 -04:00
Oleksandr Pavlyk	f14055d5cc	Change CMake's nvbench::main exported target to correspond to static library (#350 ) Previously, it corresponded to main.cu.o object file. Now it corresponds to static library libnvbench_main.a which is archive file with main.cu.o object in it. This closes #349	2026-05-13 13:10:44 -04:00
Oleksandr Pavlyk	9ea77bccaa	Implement CLI option to control warmups for cold measurements (#339 ) * Implement warmup-runs count, supported as CLI CLI option --warmup-runs implemented and documented. The warm-up counts is enforced to always be positive. This is necessary to ensure that JIT-ting has occurred, and use of blocking kernel would not result in time-outs. Test is option parser is added. * Ensure that measure_cold::run_warmup instantiates blocking kernel Because warm-up runs are executed without use of blocking kernel, the blocking kernel was not jitted until actual measurements were collected. The module loading cost incurred during the first run shows as elevated CPU time noise value for the first measurement as noted in https://github.com/NVIDIA/nvbench/pull/339 This PR adds `this->block_stream(); this->unblock_stream();` prior to executing warm-up loop with use of blocking kernel disabled. This ensures that blocking kernel is instantiated during the warm-up, but it no other kernel is launched between its launch and stream sync thus avoiding deadlocking. * Rename --warmup-runs to --cold-warmup-runs, add --cold-max-warmup-walltime Since configurable number of warmups only applies to measure_cold.cuh rename the CLI option to reflect that. Also add --cold-max-warmup-walltime (defaults to -1, i.e. disabled). If enabled, exits warmup loop before request count is reached if the wall-time expanded executign warmups exceeds this max-warmup-walltime value.	2026-05-12 14:30:08 -05:00
Oleksandr Pavlyk	ebf9f9a087	Add .coderabbit.yaml following in footsteps of CCCL (#359 )	2026-05-12 13:55:46 -05:00
Oleksandr Pavlyk	7dfbcad27c	Create directories for output files (#360 ) * QOL UX, NVBench creates directories for output JSON, MD, CSV files This closes #185 and supports specifying `--json path/to/nonexistent/folder/result.json` This would create sequence of folders where to place result.json ``` (py313) :~/repos/nvbench$ rm -rf /tmp/nested/ (py313) :~/repos/nvbench$ ./build2/bin/nvbench.example.cpp20.axes -b copy_type_and_block_size_sweep -a Type=I32 -a BlockSize=64 --jsonbin /tmp/nested/json/axes.json --md /tmp/nested/md/res.md --csv /tmp/nested/csv/res.csv > /dev/null 2>&1 (py313) :~/repos/nvbench$ tree /tmp/nested/ /tmp/nested/ ├── csv │ └── res.csv ├── json │ ├── axes.json │ ├── axes.json-bin │ │ └── 0.bin │ └── axes.json-freqs-bin │ └── 0.bin └── md └── res.md 6 directories, 5 files ``` * Add a test that non-existent output folder is created * Remove throwing custom error message. Use default * Replace static_assert(false, ...) with #error	2026-05-12 10:26:28 -05:00
Oleksandr Pavlyk	d13a0fde32	Correct cuda cccl examples per change in api (#353 )	2026-05-06 13:30:44 -05:00
Oleksandr Pavlyk	f392725015	Correct Python API signature of State.get_axis_values_as_strings (#346 ) * Correct Python API signature of State.get_axis_values_as_strings The C++ API has default boolean argument color, but Python API declared no arguments. Closes #345 * Also exercise invocation of get_axis_values_as_string with keyword argument value * Remove use of cuda.core.experimental	2026-05-04 08:40:29 -05:00
Oleksandr Pavlyk	a3364ca5c7	Port changes to the package from #323 (#337 ) Fixed relative text alignment in docstrings to fix autodoc warnigns Renamed cuda.bench.test_cpp_exception and cuda.bench.test_py_exception functions to start with underscore, signaling that these functions are internal and should not be documented Account for test_cpp_exceptions -> _test_cpp_exception, same for _py_ Make sure to reset __module__ of reexported symbols to be cuda.bench	2026-04-22 08:28:15 -05:00
Oleksandr Pavlyk	b0a46f44c2	Modularize color handling (#336 ) * Introduce function colorize to modularize colorization/no-color handling * Use sns.set_theme instead of deprecated sns.set() * Use str.format instead of legacy % syntax * Simplified iteration over list Use f-string (supported since Python 3.6) instead of str.format for better readability and performance	2026-04-14 08:09:44 -05:00
pre-commit-ci[bot]	8d23e3e73c	[pre-commit.ci] pre-commit autoupdate (#333 ) * [pre-commit.ci] pre-commit autoupdate updates: - [github.com/pre-commit/mirrors-clang-format: v21.1.8 → v22.1.2](https://github.com/pre-commit/mirrors-clang-format/compare/v21.1.8...v22.1.2) - [github.com/astral-sh/ruff-pre-commit: v0.14.10 → v0.15.9](https://github.com/astral-sh/ruff-pre-commit/compare/v0.14.10...v0.15.9) - [github.com/codespell-project/codespell: v2.4.1 → v2.4.2](https://github.com/codespell-project/codespell/compare/v2.4.1...v2.4.2) * [pre-commit.ci] auto code formatting --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>	2026-04-13 16:24:55 +00:00

1 2 3 4 5 ...

805 Commits