Repository navigation
Rebuild MOTChallenge evaluation around one fast path - #215
mikel-brostrom wants to merge 36 commits into
Conversation
bf71171 to
9aa0d03
Compare
|
This is a really big one @cheind. Please check if it aligned without your idea of the future of the package. |
|
@mikel-brostrom, thanks! Very nice features as far as i can tell from reading the description. So, essentially: it removes the generic metrics framework, all alternate solver backends, the accumulator/event-table architecture, most public modules, the CLI, and replaces them with a single specialized evaluate_motchallenge(...) API focused on MOTChallenge evaluation, correct? Are CLEAR, Identity, and HOTA the main metrics used today? I see that this PR retains the existing 18 default metrics while adding 16? more. My main concern is future extensibility: by removing MetricsHost, dynamic metric registration, and the accumulator/event representation, how difficult would it be to add a new metric family that requires different matching semantics or intermediate data not already produced by this fast path? (not that adding a new metric was easy before :)) Could you illustrate this by outlining the changes required to add one representative metric outside the current CLEAR/Identity/HOTA design? Also, besides custom metrics, are there any previously implemented non-default metrics or event-level diagnostics that users would no longer be able to compute? |
"""Compute track coverage (TCOV) as an opt-in metric family.
For each ground-truth trajectory, TCOV measures the fraction of its annotated
lifespan covered by any tracker identity linked to it through CLEAR matching.
The reported score is the mean coverage across ground-truth trajectories, so a
TCOV of 0.8 means that the tracker sees an average object for approximately 80%
of its lifetime in the evaluated frames.
Run this example with either two MOTChallenge files or two evaluation roots:
python examples/track_coverage.py path/to/gt path/to/predictions
"""
import argparse
import motmetrics as mm
class TrackCoverage(mm.MetricFamily):
"""Average per-GT-trajectory lifespan coverage by linked tracker tracks."""
name = "track_coverage"
metric_names = ("tcov",)
requirements = frozenset(("clear_statistics",))
display_names = {"tcov": "TCOV"}
formatters = {"tcov": "{:.1%}".format}
def evaluate_sequence(self, sequence, intermediates):
del sequence
per_track_coverage = intermediates.clear_statistics.track_coverage
# Preserve additive state so OVERALL can average trajectories rather
# than incorrectly averaging already-normalized sequence scores.
return float(per_track_coverage.sum()), len(per_track_coverage)
def summarize(self, partial):
coverage_sum, track_count = partial
tcov = coverage_sum / max(1, track_count)
if not 0.0 <= tcov <= 1.0:
raise ValueError("TCOV must be between 0.0 and 1.0, got {!r}".format(tcov))
return {"tcov": tcov}
def combine(self, partials):
coverage_sum = sum(partial[0] for partial in partials)
track_count = sum(partial[1] for partial in partials)
return self.summarize((coverage_sum, track_count))
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("ground_truth", help="Ground-truth file or evaluation root")
parser.add_argument("predictions", help="Prediction file or evaluation root")
parser.add_argument(
"--jobs",
type=int,
default=1,
help="Sequence worker processes (default: 1)",
)
args = parser.parse_args()
summary = mm.evaluate_motchallenge(
args.ground_truth,
args.predictions,
n_jobs=args.jobs,
extra_metric_families=TrackCoverage(),
progress=False,
)
print(summary)
if __name__ == "__main__":
main() |
LapX is now the single private assignment backend, and
The count is exactly 18 retained plus 16 added, for 34 default metrics:
Of the 34 metrics, 31 have direct TrackEval counterparts and match TrackEval 1.3.0 numerically. The remaining three: IDt, IDa, and IDm, are motmetrics-specific.
The current branch adds a smaller, explicit extension boundary through
Every family receives immutable raw ground-truth and tracker detections. It can therefore implement completely different matching semantics without changing the built-in CLEAR/HOTA matcher. |
|
I recommend that you play around with this branch before merging, as I am not fully aware of all the previous functionality and I might have missed to port something. I took the liberty to rebrand the repo a bit. Hope your are all okay with it 😅. |
|
Now also tried on dataset with distractor classes (MOT17)
|
|
@mikel-brostrom I had an in depth-look and I lilke the simplification. I will accept the PR once I'm back from vacation! Thanks! |
|
Sorry for my late response @cheind. I have also been on vacation 😅. Glad to see that this repo will get a proper update. |
|
I added some important information to the |
…unt - 2) worker limit
…2)) for arbitrary datasets
…g process_cpu_count for Python 3.13+
|
This PR is very close to something we've been building, so sharing measurements in case they're useful. Disclosure up front: I work on this at thyn-ai. We built Correctness evidence (differential suite: seeded synthetic MOTChallenge-style streams plus adversarial cases — tie-heavy integer distance matrices, all-NaN matrices, empty and one-sided frames, fragmentation boundaries,
Benchmarks (macOS arm64, Apple M4 Max; Python 3.12.14, NumPy 2.5.3, motmetrics 1.4.0; a pipeline is accumulate a whole seeded stream frame by frame, then compute the 18 MOTChallenge metrics; correctness against the oracle is asserted before timing):
i.e. 37–48x on 100–2,000-frame streams, cold and warm alike. Package shape: per-platform wheels (macOS arm64, Linux x86_64) carry the compiled kernel with the Mojo runtime vendored in, so no toolchain is needed; on every other platform the same API transparently runs on the vendored pure-Python fallback. NumPy is the only runtime dependency. Since this PR rebuilds evaluation around one fast path: if there's interest, we'd be glad to collaborate on making the kernel an official optional fast path behind the motmetrics API (e.g. an extra such as |
|
Update: the package has landed on |
Summary
This PR replaces the general-purpose accumulator and metric-host architecture with one optimized MOTChallenge evaluation path:
The new engine:
n_jobs > 1, with one parent-rendered progress row per sequence.dfis requestedWhy
Master is a broad metric framework built around accumulator event DataFrames, a dependency-resolving
MetricsHost, multiple assignment backends, and CLI applications. Its supported MOTChallenge workflow is theeval_motchallengeCLI; master does not exposeevaluate_motchallenge(...).That design adds avoidable overhead for the primary MOTChallenge use case:
motmetricseagerly imports pandas, SciPy, and supporting modulesThe new implementation focuses the package on a single supported workflow and keeps parsing, matching, aggregation, and rendering in NumPy and the standard library. Optional functionality is imported only at its use site.
Impact
API and packaging
This is intentionally a breaking, minimal API:
evaluate_motchallengeis the only public exporthota_alphas, distance threshold, metric selection, worker count, and progress remain configurable through this entrypointmetricsselects fewer CLEAR or Identity fieldsprint(summary)does not require pandassummary.dfrequirespip install "motmetrics[dataframe]"; dependencies are never installed at runtimemotmetrics.metrics,motmetrics.mot,motmetrics.io,motmetrics.lap, andmotmetrics.utilsare intentionally absentMOTAccumulator,MetricsHost, custom registration, solver selection, CLI metric applications, aliases, deprecated APIs, and compatibility shims are intentionally absent__version__is removed; useimportlib.metadata.version("motmetrics")Required Python 3 dependencies decrease from four to two:
pandas is available as the optional
dataframeextra.Assignment backend: SciPy to LapX
The previous solver layer supported SciPy, lapsolver, OR-Tools, and Munkres, discovered available packages at runtime, and dispatched each assignment through a generic public solver API. SciPy was both a required dependency and the normal fallback.
The new engine uses one private backend:
lap.lapjvfrom LapX. The same backend handles frame-level CLEAR matching, global Identity assignment, and HOTA matching. Rectangular matrices use LapX's extended-cost mode. For forbiddenNaNor infinite edges, the wrapper substitutes a finite cost that cannot improve a valid solution, solves the dense assignment, and then removes any substituted pairs. This preserves the previous missing-edge semantics without backend-specific branches.LapX was selected for the complete evaluation path, not because every isolated assignment is necessarily faster than SciPy:
LapX versus SciPy timing
The backend comparison was rerun on the same seven MOT17 ablation sequences using macOS arm64, Python 3.14.5, NumPy 2.5.1, LapX 0.9.4, and SciPy 1.18.0. Backend order was alternated to reduce ordering and cache bias.
Fresh import timing includes Python startup, NumPy, and the assignment backend:
SciPy adds 0.427 s to the median fresh import and takes 4.80x as long to import.
Warm evaluation excludes interpreter and backend import time but includes MOT17 file loading, IoU preparation, all metrics, aggregation, and rendering. Each backend was warmed once before eleven measured alternating runs:
Once both backends are already imported, SciPy is 0.022 s, or approximately 7.5%, faster on this workload. This is why the backend choice is not presented as a per-assignment speed win.
Fresh end-to-end evaluation includes interpreter startup, all imports, MOT17 parsing, IoU preparation, all metrics, aggregation, and rendering:
The SciPy import penalty dominates its small warm-compute advantage. The LapX path is 2.04x faster end to end and saves 0.413 s per fresh MOT17 process in this comparison.
The benchmark changed only the linear-assignment function behind the current evaluator. Results from both backends agree to a maximum absolute difference of
2.220e-16, which is floating-point rounding noise.Assignment correctness is covered by the MOT17 comparison below: every metric with a TrackEval counterpart matches across all seven sequences and
OVERALL, with only floating-point differences around1e-16to1e-15.Metrics added relative to master
Master reports 18 default metrics. This PR retains those metrics and adds 16, producing 34 default columns:
MTRPTRMLRMODAsMOTACLR_F1FP/FrameHOTADetAAssADetReDetPrAssReAssPrLocAOWTAMT/PT/ML classification, fragmentation handling, identity switches, and HOTA aggregation follow TrackEval semantics.
The built wheel decreases from 160,959 to 27,559 bytes, an 82.9% reduction.
Performance
Performance is compared directly with TrackEval 1.3.0. Master timings are intentionally excluded.
End-to-end MOT17 benchmark
The benchmark used:
n_jobs=1for the current evaluator, which is the fastest fresh-process setting for this workloadEach timed process includes:
OVERALLaggregationTrackEval / current median runtime ratio: 2.73x. The current implementation completes the full end-to-end evaluation in approximately 36.7% of TrackEval's time.
Backend order was alternated on every run to reduce ordering and cache bias. Both implementations received the same source files and evaluation threshold. The current evaluator rendered its complete 34-column result; the TrackEval path rendered the 31 directly comparable public values, including
Count.GT_IDs.The TrackEval measurement uses its metric API directly instead of the higher-overhead dataset evaluator and CLI, making the comparison conservative in TrackEval's favor. Metric correctness is reported separately below and confirms numerical parity across all seven sequences and
OVERALL.Test
MOT17 metric parity with TrackEval 1.3.0
Metric values were compared only with TrackEval, not with master.
The current public evaluator and TrackEval 1.3.0 received identical frame, identity, bounding-box, and IoU data from these seven MOT17 ablation sequences:
MOT17-02-FRCNNMOT17-04-FRCNNMOT17-05-FRCNNMOT17-09-FRCNNMOT17-10-FRCNNMOT17-11-FRCNNMOT17-13-FRCNNAll seven sequence rows and the combined
OVERALLrow were compared. “Max abs difference” below is the largest difference across all eight rows.Results:
OVERALL7.772e-16, forLocA1.887e-15, atMOT17-04-FRCNN,LocA, alpha0.75IDt,IDa, andIDmare motmetrics-specific event diagnostics without TrackEval 1.3.0 counterpartsTrackEval reports MOTP as similarity, while the public motmetrics table reports distance. TrackEval MOTP was converted with
1 - MOTPbefore comparison.GTwas compared with TrackEvalCount.GT_IDs.The non-zero differences are approximately
1e-16to1e-15and are floating-point rounding noise.Additional validation
OVERALLtwine checkgit diff --checkpasses