Analyzing measurements¶
The iperf3_lib.analysis module provides standard-library calculations
over canonical results. It does not run a test, load libiperf, export metrics, or
reparse native JSON. Comparison checks resolve declared evidence pointers to
confirm receipts exist; they do not infer values from native fields.
A completed test and adequate measurement data are separate requirements. Every
analysis returns quality: complete, partial, or insufficient_data, plus
canonical evidence pointers and diagnostics. These describe data and methodology.
They do not identify physical bottlenecks or decide a performance acceptance rule.
Measured throughput and interval stability¶
from iperf3_lib.analysis import Selection, interval_stability, summary_throughput
received = summary_throughput(
result, direction="client_to_server", observation="receiver"
)
print(received.quality, received.throughput_bps)
stability = interval_stability(
result,
selection=Selection("client_to_server", "sender", scope="aggregate"),
threshold_bps=20_000_000,
quantiles=(0.5, 0.95),
)
print(stability.coverage, stability.diagnostics)
Choose an observation that the result actually contains. A forward client's
intervals commonly describe the sender; the reverse client's intervals commonly
describe the receiver. Terminal flow summaries may contain both observations.
Direction, observer and aggregate/component-stream scope are independent choices.
Selection(..., scope="stream", stream_id=5) selects one run-local socket.
summary_throughput calculates 8 * bytes / measured_duration_seconds for the
chosen flow endpoint. It never substitutes requested test duration, wrapper
operation time, a reported bitrate, or summary_mbps. Missing bytes or a zero
measured duration produces insufficient data; measured zero bytes remains zero.
Failed or incomplete executions cannot supply successful performance analysis.
For intervals with durations d[i] and rates r[i]:
| Output | Definition |
|---|---|
| Minimum interval-average rate | min(r) |
| Duration-weighted mean | sum(d * r) / sum(d) |
| Duration-weighted population standard deviation | sqrt(sum(d * (r - mean)**2) / sum(d)) |
| Coefficient of variation | Standard deviation divided by mean; unavailable when mean is zero |
| Duration-weighted interval-average quantile | Smallest rate whose cumulative duration reaches the requested fraction, sorting by rate; no interpolation |
| Fraction of measured time below threshold | sum(d where r < threshold) / sum(d); strict less-than |
| Interval bytes throughput | 8 * sum(bytes) / sum(d) when every included interval has bytes |
A one-second interval at 8 Mbps and a three-second interval at 2 Mbps have a 3.5 Mbps weighted mean, a 2 Mbps weighted median, and 75% of measured time below 4 Mbps. The standard deviation is approximately 2.598 Mbps. These are statistics of interval-average throughput, not packet throughput or latency percentiles.
Omissions, missing values and coverage¶
Explicitly omitted warm-up intervals are excluded before overlap detection;
native warm-up clocks can overlap measured intervals. IntervalPolicy() excludes
unknown omission states, requires two intervals for stability, and uses a
one-microsecond boundary tolerance. Applications can explicitly choose
unknown_omission="include" or derive_duration_from_bounds=True; these decisions
are retained and the analysis becomes partial.
Bytes and measured duration determine each rate when available. A reported rate can support partial stability when bytes are absent, but cannot establish complete bytes throughput. A disagreement between bytes/time and reported rate is diagnosed. Missing durations, boundaries and provenance are excluded with counts. Gaps reduce coverage without inserting zeros. Overlapping or duplicate selected intervals make stability insufficient. Aggregate and per-stream observations are never added together. Unknown coverage denominators stay unavailable.
coverage records selected, included, omitted, unknown-omission and invalid/missing
counts, included measured seconds, byte-covered seconds, observed span and measured
time fraction. Insufficient stability can still retain descriptive statistics for
a single interval; callers must inspect quality before treating them as adequate
stability evidence.
Stream balance and stream-count experiments¶
from iperf3_lib.analysis import AnalysisTrial, ComparisonPolicy, stream_balance, stream_scaling
balance = stream_balance(result, direction="client_to_server", observation="receiver")
scaling = stream_scaling(
[AnalysisTrial("one-stream", one_stream_result),
AnalysisTrial("four-streams", four_stream_result)],
direction="client_to_server",
observation="receiver",
compatibility=ComparisonPolicy("lab-run", ("client-host", "server-host")),
best_fraction=0.95,
minimum_valid_trials=1,
)
print(scaling.smallest_tested_qualifying_count, scaling.diagnostics)
Each stream uses its own endpoint byte count and measured duration. The report
includes all selected streams, valid/expected counts, coverage, minimum/maximum
rate ratio, population rate CV and Jain's fairness index
sum(r)**2 / (N * sum(r**2)). Missing streams remain missing; all-zero rates make
relative balance statistics unavailable. Socket IDs only identify streams within
one native run.
For mixed UDP terminal summaries whose endpoint attribution is unavailable,
source="intervals" explicitly selects attributable per-stream intervals. It
does not relabel the mixed end observations. The supplied interval policy also
applies to this source.
Scaling groups valid trials by verified parallel stream count and uses the median
measured flow throughput in each group. It chooses the smallest tested count
with median at least best_fraction * best_observed_median, subject to the
minimum-valid-trials requirement. Failed trials, missing evidence and insufficient
repetitions remain visible. All-zero medians yield no recommendation. There is no
interpolation of untested counts, confidence interval, or physical optimum claim.
Reusing compatibility checks¶
from iperf3_lib.analysis import check_compatibility
comparison = check_compatibility(
[AnalysisTrial("current", current), AnalysisTrial("baseline", baseline)],
policy=ComparisonPolicy("regression-check", ("client-host", "server-host")),
)
print(comparison.compatible, comparison.fingerprints, comparison.diagnostics)
ComparisonPolicy records caller-owned group and endpoint identities. Concrete
fingerprints compare protocol, execution method, native version/system information,
server/port, configured duration/omit, parallel streams, native rate, block size,
TOS, and rate intent. Verified settings require an existing receipt under raw or
a namespaced extensions object. Request values and successful setter calls are
insufficient. Native environment strings include host identity; deliberate changes
need an explicit policy reason.
varying_fields=("parallel",) declares a measured experimental variable: its
values must be known, but equality is not required. allowed_differences maps
field names to nonempty reasons, for example {"native_version": "Intentional
version comparison"}. An allowed observed difference makes quality partial and
retains its reason. Neither mechanism supplies missing evidence. Baseline checks
keep parallel count fixed by default; scaling automatically varies parallel,
and sequential asymmetry automatically varies execution method.
When the rate-intent extension records the same aggregate target per direction,
per-stream rates may differ only if every verified rate equals
aggregate_target // verified_parallel. Fixed per-stream and fixed aggregate
experiments retain different intent fingerprints. Compatibility alone does not
establish execution success or adequate throughput measurements.
Advanced configuration in comparisons¶
An advanced ClientConfig field becomes part of every trial's comparison
fingerprint when any retained request gives it a nondefault value, or when
ComparisonPolicy.varying_fields or allowed_differences explicitly names it.
This includes mptcp and json_stream. Every compared trial then needs verified
native observations for that field, including trials that requested its default.
The request selects what to compare; the returned native value or matching
getter receipt supplies the observation.
For example, comparing a no_delay=True trial with a default-configured trial
requires retained evidence of both native no-delay values. An absent receipt
for the default-configured peer cannot be replaced with an assumed False.
Equal requests also cannot fill missing observations. A known difference must
be a declared varying field or have a reason in allowed_differences; neither
policy permits unknown values.
Options accepted by the native parser may lack a returned value or exposed getter. Those runs remain available for inspection, but comparisons requiring that observation are incompatible. Request validity, execution success and comparison eligibility are separate conclusions. These rules apply equally to retained baselines and advanced settings fixed across sweep cells.
Default-only comparisons retain their existing fingerprints. Archived v1
reports validate the recorded advanced fields with frozen compatibility rules;
reading an old report does not add fields from a newer ClientConfig.
Directional asymmetry¶
simultaneous_asymmetry(result, observation="receiver") requires both directions
from one explicitly bidirectional run. sequential_asymmetry(forward, reverse,
observation="receiver", methodology=...) requires separate forward and reverse
runs and a SequentialMethodology(pair_id, comparison_policy, execution_order,
cooldown_seconds). Execution order is forward_then_reverse or
reverse_then_forward. Unknown cooldown is explicit None and makes quality
partial. Caller/runner provenance establishes order; wall clocks alone do not.
Both APIs use the same chosen observation for both directions. With measured
rates F and R, they report F - R, F / R when R > 0, and
(F - R) / max(F, R) when either is positive. Undefined or out-of-range ratios
remain unavailable. Sequential and simultaneous methodologies remain distinct in
the output. Neither compares a sender observation to a receiver observation.
TCP and endpoint CPU evidence¶
IntervalStats.tcp, per-stream sender SumStats.tcp, and Result.cpu preserve
qualified evidence independently of analysis. RTT fields are seconds; congestion
window, advertised send window and path MTU are bytes. TCP interval RTT is a
smoothed local TCP sample. Native min/max/mean RTT summaries describe sampled TCP
information, not packet latency percentiles. Source-defined -1 getter values
become unavailable with explicit availability evidence. See the native
TCP getter units
and summary sampling.
Normalization is qualified against Linux output from libiperf 3.19.1 and 3.21. TCP information belongs to the local socket. Only attributable local sender observations are normalized; remote sender entries can contain local receiver TCP values or placeholders. Unknown producers, platforms or contradictory sender provenance retain raw evidence and explicit uncertainty. The retained fixtures cover both versions and both reporting endpoints.
CPU values describe the iperf process, retaining endpoint identity and local/remote
provenance. Total/user/system percentages can exceed 100% because process CPU
time can accumulate across threads. The native implementation divides process CPU
time by elapsed time (source).
Qualified local CPU values are retained. Final remote CPU is established in client
output; server output's remote fields are kept unknown. Unestablished endpoint
identity remains "unknown". Each present measurement has an evidence pointer.
TCP retransmissions are not an exact packet-loss percentage. UDP jitter is not application latency. CPU, RTT and window observations are facts with provenance; these APIs produce no speculative root-cause hypotheses.