Skip to content

Rebuild MOTChallenge evaluation around one fast path - #215

Open
mikel-brostrom wants to merge 36 commits into
cheind:developfrom
mikel-brostrom:agent/speed-up-motchallenge-eval
Open

mikel-brostrom wants to merge 36 commits into
cheind:developfrom
mikel-brostrom:agent/speed-up-motchallenge-eval

Conversation

@mikel-brostrom

@mikel-brostrom mikel-brostrom commented Jul 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR replaces the general-purpose accumulator and metric-host architecture with one optimized MOTChallenge evaluation path:

motmetrics.evaluate_motchallenge(...)

The new engine:

  • computes CLEAR, Identity, and HOTA metrics in one canonical path
  • always calculates HOTA; there is no alternate non-HOTA execution branch
  • returns 34 metrics by default, compared with 18 on master
  • parses inputs directly into compact NumPy-backed sequence data
  • reuses matching, identity, and HOTA state instead of constructing pandas event tables
  • uses one process per sequence when n_jobs > 1, with one parent-rendered progress row per sequence
  • returns a lightweight native summary; pandas is imported only when .df is requested
  • uses NumPy and LapX as its only required runtime dependencies
  • supports CPython 3.8 through 3.14
  • removes legacy computation paths, CLI applications, compatibility shims, deprecated entrypoints, dynamic metric registration, and solver dispatch

Why

Master is a broad metric framework built around accumulator event DataFrames, a dependency-resolving MetricsHost, multiple assignment backends, and CLI applications. Its supported MOTChallenge workflow is the eval_motchallenge CLI; master does not expose evaluate_motchallenge(...).

That design adds avoidable overhead for the primary MOTChallenge use case:

  • importing motmetrics eagerly imports pandas, SciPy, and supporting modules
  • every sequence constructs and processes pandas event tables
  • metric calculation passes through generic registration and dependency machinery
  • parsing, matching, and aggregation are split across several public workflows
  • HOTA and its diagnostic metrics are not available

The new implementation focuses the package on a single supported workflow and keeps parsing, matching, aggregation, and rendering in NumPy and the standard library. Optional functionality is imported only at its use site.

Impact

API and packaging

This is intentionally a breaking, minimal API:

import motmetrics as mm

summary = mm.evaluate_motchallenge(groundtruths, predictions)
print(summary)
  • evaluate_motchallenge is the only public export
  • hota_alphas, distance threshold, metric selection, worker count, and progress remain configurable through this entrypoint
  • HOTA is always calculated, including when metrics selects fewer CLEAR or Identity fields
  • print(summary) does not require pandas
  • summary.df requires pip install "motmetrics[dataframe]"; dependencies are never installed at runtime
  • legacy imports such as motmetrics.metrics, motmetrics.mot, motmetrics.io, motmetrics.lap, and motmetrics.utils are intentionally absent
  • MOTAccumulator, MetricsHost, custom registration, solver selection, CLI metric applications, aliases, deprecated APIs, and compatibility shims are intentionally absent
  • package-level __version__ is removed; use importlib.metadata.version("motmetrics")
  • UA-DETRAC MAT input is removed; DETRAC XML remains supported
  • the wheel contains only eight runtime Python modules and package metadata

Required Python 3 dependencies decrease from four to two:

Master Current
NumPy NumPy
pandas LapX
SciPy
xmltodict

pandas is available as the optional dataframe extra.

Assignment backend: SciPy to LapX

The previous solver layer supported SciPy, lapsolver, OR-Tools, and Munkres, discovered available packages at runtime, and dispatched each assignment through a generic public solver API. SciPy was both a required dependency and the normal fallback.

The new engine uses one private backend: lap.lapjv from LapX. The same backend handles frame-level CLEAR matching, global Identity assignment, and HOTA matching. Rectangular matrices use LapX's extended-cost mode. For forbidden NaN or infinite edges, the wrapper substitutes a finite cost that cannot improve a valid solution, solves the dense assignment, and then removes any substituted pairs. This preserves the previous missing-edge semantics without backend-specific branches.

LapX was selected for the complete evaluation path, not because every isolated assignment is necessarily faster than SciPy:

  • it removes SciPy and its import cost from the required dependency set
  • it removes runtime solver discovery, string dispatch, optional fallbacks, and solver-specific normalization
  • it gives every metric one consistent assignment implementation
  • the sparse connected-component work formerly delegated to SciPy is handled internally, so the rest of SciPy is no longer needed
  • it keeps cold-start and fresh-process evaluation lightweight, which matters more here than a small difference across thousands of already-warm, tiny assignments

LapX versus SciPy timing

The backend comparison was rerun on the same seven MOT17 ablation sequences using macOS arm64, Python 3.14.5, NumPy 2.5.1, LapX 0.9.4, and SciPy 1.18.0. Backend order was alternated to reduce ordering and cache bias.

Fresh import timing includes Python startup, NumPy, and the assignment backend:

Backend Median Minimum Maximum Fresh processes
LapX 0.113 s 0.073 s 0.518 s 15
SciPy 0.540 s 0.425 s 0.745 s 15

SciPy adds 0.427 s to the median fresh import and takes 4.80x as long to import.

Warm evaluation excludes interpreter and backend import time but includes MOT17 file loading, IoU preparation, all metrics, aggregation, and rendering. Each backend was warmed once before eleven measured alternating runs:

Backend Median Minimum Maximum Runs
LapX 0.293 s 0.280 s 0.299 s 11
SciPy 0.271 s 0.261 s 0.282 s 11

Once both backends are already imported, SciPy is 0.022 s, or approximately 7.5%, faster on this workload. This is why the backend choice is not presented as a per-assignment speed win.

Fresh end-to-end evaluation includes interpreter startup, all imports, MOT17 parsing, IoU preparation, all metrics, aggregation, and rendering:

Backend Median Minimum Maximum Fresh processes
LapX 0.396 s 0.371 s 0.460 s 9
SciPy 0.809 s 0.724 s 0.863 s 9

The SciPy import penalty dominates its small warm-compute advantage. The LapX path is 2.04x faster end to end and saves 0.413 s per fresh MOT17 process in this comparison.

The benchmark changed only the linear-assignment function behind the current evaluator. Results from both backends agree to a maximum absolute difference of 2.220e-16, which is floating-point rounding noise.

Assignment correctness is covered by the MOT17 comparison below: every metric with a TrackEval counterpart matches across all seven sequences and OVERALL, with only floating-point differences around 1e-16 to 1e-15.

Metrics added relative to master

Master reports 18 default metrics. This PR retains those metrics and adds 16, producing 34 default columns:

Added metric Family Meaning
MTR CLEAR Mostly-tracked trajectories divided by ground-truth trajectories
PTR CLEAR Partially-tracked trajectories divided by ground-truth trajectories
MLR CLEAR Mostly-lost trajectories divided by ground-truth trajectories
MODA CLEAR Detection accuracy without the identity-switch penalty
sMOTA CLEAR MOTA variant incorporating localization similarity
CLR_F1 CLEAR Harmonic mean of CLEAR precision and recall
FP/Frame CLEAR False-positive detections per evaluated frame
HOTA HOTA Geometric mean of detection and association accuracy
DetA HOTA HOTA detection accuracy
AssA HOTA HOTA association accuracy
DetRe HOTA HOTA detection recall
DetPr HOTA HOTA detection precision
AssRe HOTA HOTA association recall
AssPr HOTA HOTA association precision
LocA HOTA Mean localization similarity of HOTA matches
OWTA HOTA Open-world tracking accuracy

MT/PT/ML classification, fragmentation handling, identity switches, and HOTA aggregation follow TrackEval semantics.

The built wheel decreases from 160,959 to 27,559 bytes, an 82.9% reduction.

Performance

Performance is compared directly with TrackEval 1.3.0. Master timings are intentionally excluded.

End-to-end MOT17 benchmark

The benchmark used:

  • the same seven MOT17 ablation sequences used for metric parity
  • macOS arm64
  • Python 3.14.5
  • TrackEval 1.3.0
  • nine alternating fresh processes per implementation
  • n_jobs=1 for the current evaluator, which is the fastest fresh-process setting for this workload
  • every sample retained; no warm-up sample was discarded

Each timed process includes:

  1. Python interpreter startup
  2. all package imports
  3. reading and parsing all ground-truth and tracker files
  4. ground-truth filtering and frame/identity preparation
  5. IoU matrix construction
  6. all CLEAR, Identity, Count, and 19-threshold HOTA calculations
  7. sequence and OVERALL aggregation
  8. result rendering
Implementation Median Minimum Maximum Fresh processes
Current 0.422 s 0.363 s 0.490 s 9
TrackEval 1.3.0 1.150 s 1.081 s 1.358 s 9

TrackEval / current median runtime ratio: 2.73x. The current implementation completes the full end-to-end evaluation in approximately 36.7% of TrackEval's time.

Backend order was alternated on every run to reduce ordering and cache bias. Both implementations received the same source files and evaluation threshold. The current evaluator rendered its complete 34-column result; the TrackEval path rendered the 31 directly comparable public values, including Count.GT_IDs.

The TrackEval measurement uses its metric API directly instead of the higher-overhead dataset evaluator and CLI, making the comparison conservative in TrackEval's favor. Metric correctness is reported separately below and confirms numerical parity across all seven sequences and OVERALL.

Test

MOT17 metric parity with TrackEval 1.3.0

Metric values were compared only with TrackEval, not with master.

The current public evaluator and TrackEval 1.3.0 received identical frame, identity, bounding-box, and IoU data from these seven MOT17 ablation sequences:

  • MOT17-02-FRCNN
  • MOT17-04-FRCNN
  • MOT17-05-FRCNN
  • MOT17-09-FRCNN
  • MOT17-10-FRCNN
  • MOT17-11-FRCNN
  • MOT17-13-FRCNN

All seven sequence rows and the combined OVERALL row were compared. “Max abs difference” below is the largest difference across all eight rows.

Metric Current OVERALL TrackEval OVERALL Signed difference Max abs difference
IDF1 0.825859619783 0.825859619783 0 0
IDP 0.903414225573 0.903414225573 0 0
IDR 0.760567823344 0.760567823344 0 0
Rcll 0.795249582483 0.795249582483 0 0
Prcn 0.944609755560 0.944609755560 0 0
GT 339 339 0 0
MT 174 174 0 0
PT 121 121 0 0
ML 44 44 0 0
MTR 0.513274336283 0.513274336283 0 0
PTR 0.356932153392 0.356932153392 0 0
MLR 0.129793510324 0.129793510324 0 0
FP 2513 2513 0 0
FN 11034 11034 0 0
IDs 162 162 0 0
FM 577 577 0 0
MOTA 0.745611430692 0.745611430692 0 1.110e-16
MODA 0.748617554277 0.748617554277 0 0
MOTP 0.145449082739 0.145449082739 -3.053e-16 4.718e-16
sMOTA 0.629943108372 0.629943108372 3.331e-16 4.441e-16
CLR_F1 0.863518673370 0.863518673370 0 0
FP/Frame 0.947586726998 0.947586726998 0 0
IDt 76 N/A N/A N/A
IDa 80 N/A N/A N/A
IDm 21 N/A N/A N/A
HOTA 0.688242604918 0.688242604918 1.110e-16 1.110e-16
DetA 0.645164441753 0.645164441753 0 0
AssA 0.738271507003 0.738271507003 3.331e-16 3.331e-16
DetRe 0.697771288492 0.697771288492 0 0
DetPr 0.828823530094 0.828823530094 0 0
AssRe 0.774636914924 0.774636914924 1.110e-16 5.551e-16
AssPr 0.872836864001 0.872836864001 0 2.220e-16
LocA 0.871850812219 0.871850812219 6.661e-16 7.772e-16
OWTA 0.717276213217 0.717276213217 1.110e-16 2.220e-16

Results:

  • 34 current public metrics checked
  • 31 metrics have one-to-one TrackEval equivalents
  • all 31 match across every sequence and OVERALL
  • 22 metrics match exactly in every row
  • largest public-value difference: 7.772e-16, for LocA
  • largest difference across every individual HOTA threshold: 1.887e-15, at MOT17-04-FRCNN, LocA, alpha 0.75
  • IDt, IDa, and IDm are motmetrics-specific event diagnostics without TrackEval 1.3.0 counterparts

TrackEval reports MOTP as similarity, while the public motmetrics table reports distance. TrackEval MOTP was converted with 1 - MOTP before comparison. GT was compared with TrackEval Count.GT_IDs.

The non-zero differences are approximately 1e-16 to 1e-15 and are floating-point rounding noise.

Additional validation

  • current suite with TrackEval 1.3.0 installed: 50 passed
  • exact TrackEval parity tests cover all 34 public fields on the bundled TUD sequences and OVERALL
  • clean CPython 3.8.20 environment: 49 passed, 1 optional TrackEval test skipped
  • serial and parallel MOT17 evaluations produce identical summaries
  • master baseline suite in its isolated environment: 48 passed
  • wheel and source distribution build successfully
  • package metadata passes twine check
  • Ruff validation passes
  • git diff --check passes
  • package audit confirms one public entrypoint, no HOTA opt-out, no registration host, no deprecated metric path, and no compatibility shims
  • wheel audit confirms no CLI entrypoints, tests, fixtures, SciPy dependency, or legacy public modules

@mikel-brostrom
mikel-brostrom marked this pull request as draft July 17, 2026 09:41
@mikel-brostrom
mikel-brostrom force-pushed the agent/speed-up-motchallenge-eval branch from bf71171 to 9aa0d03 Compare July 17, 2026 19:44
@mikel-brostrom mikel-brostrom changed the title Speed up high-level MOTChallenge evaluation Rebuild MOTChallenge evaluation around one fast path Jul 17, 2026
@mikel-brostrom

Copy link
Copy Markdown
Contributor Author

This is a really big one @cheind. Please check if it aligned without your idea of the future of the package.

@cheind

cheind commented Jul 18, 2026

Copy link
Copy Markdown
Owner

@mikel-brostrom, thanks! Very nice features as far as i can tell from reading the description.

So, essentially: it removes the generic metrics framework, all alternate solver backends, the accumulator/event-table architecture, most public modules, the CLI, and replaces them with a single specialized evaluate_motchallenge(...) API focused on MOTChallenge evaluation, correct?

Are CLEAR, Identity, and HOTA the main metrics used today? I see that this PR retains the existing 18 default metrics while adding 16? more. My main concern is future extensibility: by removing MetricsHost, dynamic metric registration, and the accumulator/event representation, how difficult would it be to add a new metric family that requires different matching semantics or intermediate data not already produced by this fast path? (not that adding a new metric was easy before :))

Could you illustrate this by outlining the changes required to add one representative metric outside the current CLEAR/Identity/HOTA design? Also, besides custom metrics, are there any previously implemented non-default metrics or event-level diagnostics that users would no longer be able to compute?

@mikel-brostrom

mikel-brostrom commented Jul 18, 2026 •

Copy link
Copy Markdown
Contributor Author

@cheind

"""Compute track coverage (TCOV) as an opt-in metric family.

For each ground-truth trajectory, TCOV measures the fraction of its annotated
lifespan covered by any tracker identity linked to it through CLEAR matching.
The reported score is the mean coverage across ground-truth trajectories, so a
TCOV of 0.8 means that the tracker sees an average object for approximately 80%
of its lifetime in the evaluated frames.

Run this example with either two MOTChallenge files or two evaluation roots:

    python examples/track_coverage.py path/to/gt path/to/predictions
"""

import argparse

import motmetrics as mm


class TrackCoverage(mm.MetricFamily):
    """Average per-GT-trajectory lifespan coverage by linked tracker tracks."""

    name = "track_coverage"
    metric_names = ("tcov",)
    requirements = frozenset(("clear_statistics",))
    display_names = {"tcov": "TCOV"}
    formatters = {"tcov": "{:.1%}".format}

    def evaluate_sequence(self, sequence, intermediates):
        del sequence
        per_track_coverage = intermediates.clear_statistics.track_coverage

        # Preserve additive state so OVERALL can average trajectories rather
        # than incorrectly averaging already-normalized sequence scores.
        return float(per_track_coverage.sum()), len(per_track_coverage)

    def summarize(self, partial):
        coverage_sum, track_count = partial
        tcov = coverage_sum / max(1, track_count)
        if not 0.0 <= tcov <= 1.0:
            raise ValueError("TCOV must be between 0.0 and 1.0, got {!r}".format(tcov))
        return {"tcov": tcov}

    def combine(self, partials):
        coverage_sum = sum(partial[0] for partial in partials)
        track_count = sum(partial[1] for partial in partials)
        return self.summarize((coverage_sum, track_count))


def main():
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("ground_truth", help="Ground-truth file or evaluation root")
    parser.add_argument("predictions", help="Prediction file or evaluation root")
    parser.add_argument(
        "--jobs",
        type=int,
        default=1,
        help="Sequence worker processes (default: 1)",
    )
    args = parser.parse_args()

    summary = mm.evaluate_motchallenge(
        args.ground_truth,
        args.predictions,
        n_jobs=args.jobs,
        extra_metric_families=TrackCoverage(),
        progress=False,
    )
    print(summary)


if __name__ == "__main__":
    main()

@mikel-brostrom
mikel-brostrom marked this pull request as ready for review July 18, 2026 20:39
@mikel-brostrom

Copy link
Copy Markdown
Contributor Author

it removes the generic metrics framework, all alternate solver backends, the accumulator/event-table architecture, most public modules, the CLI, and replaces them with a single specialized evaluate_motchallenge(...) API focused on MOTChallenge evaluation, correct?

LapX is now the single private assignment backend, and evaluate_motchallenge(...) is the only evaluation entrypoint.
It does not remove accumulator or event concepts completely. The optimized path uses a private, compact accumulator internally instead of pandas event tables. It also exposes MetricFamily for explicit per-evaluation extensions.

Are CLEAR, Identity, and HOTA the main metrics used today? I see that this PR retains the existing 18 default metrics while adding 16?

The count is exactly 18 retained plus 16 added, for 34 default metrics:

  • 7 CLEAR derivatives: MTR, PTR, MLR, MODA, sMOTA, CLR_F1, and FP/Frame.
  • 9 HOTA metrics: HOTA, DetA, AssA, DetRe, DetPr, AssRe, AssPr, LocA, and OWTA.

Of the 34 metrics, 31 have direct TrackEval counterparts and match TrackEval 1.3.0 numerically. The remaining three: IDt, IDa, and IDm, are motmetrics-specific.

My main concern is future extensibility: by removing MetricsHost, dynamic metric registration, and the accumulator/event representation, how difficult would it be to add a new metric family that requires different matching semantics or intermediate data not already produced by this fast path? (not that adding a new metric was easy before :))

The current branch adds a smaller, explicit extension boundary through MetricFamily. It implements three operations:

  • evaluate_sequence(...) produces compact per-sequence state.
  • summarize(...) produces the sequence metrics.
  • combine(...) combines the original states correctly for OVERALL.

Every family receives immutable raw ground-truth and tracker detections. It can therefore implement completely different matching semantics without changing the built-in CLEAR/HOTA matcher.

#215 (comment)

@mikel-brostrom

mikel-brostrom commented Jul 18, 2026 •

Copy link
Copy Markdown
Contributor Author

I recommend that you play around with this branch before merging, as I am not fully aware of all the previous functionality and I might have missed to port something. I took the liberty to rebrand the repo a bit. Hope your are all okay with it 😅.

@mikel-brostrom
mikel-brostrom marked this pull request as draft July 19, 2026 08:29
@mikel-brostrom

Copy link
Copy Markdown
Contributor Author

Now also tried on dataset with distractor classes (MOT17)

Sequence Metric MOTMetrics TrackEval Difference Status
MOT17-02-FRCNN idf1 0.124822190612 0.124822190612 0.000e+00 PASS
MOT17-02-FRCNN idp 0.513157894737 0.513157894737 0.000e+00 PASS
MOT17-02-FRCNN idr 0.0710526315789 0.0710526315789 0.000e+00 PASS
MOT17-02-FRCNN recall 0.133198380567 0.133198380567 0.000e+00 PASS
MOT17-02-FRCNN precision 0.961988304094 0.961988304094 0.000e+00 PASS
MOT17-02-FRCNN num_unique_objects 53 53 0.000e+00 PASS
MOT17-02-FRCNN mostly_tracked 0 0 0.000e+00 PASS
MOT17-02-FRCNN partially_tracked 8 8 0.000e+00 PASS
MOT17-02-FRCNN mostly_lost 45 45 0.000e+00 PASS
MOT17-02-FRCNN mtr 0 0 0.000e+00 PASS
MOT17-02-FRCNN ptr 0.150943396226 0.150943396226 0.000e+00 PASS
MOT17-02-FRCNN mlr 0.849056603774 0.849056603774 0.000e+00 PASS
MOT17-02-FRCNN num_false_positives 52 52 0.000e+00 PASS
MOT17-02-FRCNN num_misses 8564 8564 0.000e+00 PASS
MOT17-02-FRCNN num_switches 239 239 0.000e+00 PASS
MOT17-02-FRCNN num_fragmentations 276 276 0.000e+00 PASS
MOT17-02-FRCNN mota 0.103744939271 0.103744939271 0.000e+00 PASS
MOT17-02-FRCNN moda 0.127935222672 0.127935222672 0.000e+00 PASS
MOT17-02-FRCNN motp 0.143751241631 0.143751241631 1.388e-16 PASS
MOT17-02-FRCNN smota 0.0845975066816 0.0845975066816 1.388e-17 PASS
MOT17-02-FRCNN clr_f1 0.23399715505 0.23399715505 0.000e+00 PASS
MOT17-02-FRCNN fp_per_frame 0.173913043478 0.173913043478 0.000e+00 PASS
MOT17-02-FRCNN hota 0.125835737499 0.125835737499 2.776e-17 PASS
MOT17-02-FRCNN deta 0.116294513568 0.116294513568 0.000e+00 PASS
MOT17-02-FRCNN assa 0.140171458969 0.140171458969 8.327e-17 PASS
MOT17-02-FRCNN detre 0.117616663115 0.117616663115 0.000e+00 PASS
MOT17-02-FRCNN detpr 0.849453678055 0.849453678055 0.000e+00 PASS
MOT17-02-FRCNN assre 0.142164241503 0.142164241503 1.110e-16 PASS
MOT17-02-FRCNN asspr 0.935906499759 0.935906499759 0.000e+00 PASS
MOT17-02-FRCNN loca 0.868871025922 0.868871025922 1.110e-16 PASS
MOT17-02-FRCNN owta 0.126737049264 0.126737049264 2.776e-17 PASS
MOT17-04-FRCNN idf1 0.506581141758 0.506581141758 0.000e+00 PASS
MOT17-04-FRCNN idp 0.667162262621 0.667162262621 0.000e+00 PASS
MOT17-04-FRCNN idr 0.408305070725 0.408305070725 0.000e+00 PASS
MOT17-04-FRCNN recall 0.600876830176 0.600876830176 0.000e+00 PASS
MOT17-04-FRCNN precision 0.981820639319 0.981820639319 0.000e+00 PASS
MOT17-04-FRCNN num_unique_objects 69 69 0.000e+00 PASS
MOT17-04-FRCNN mostly_tracked 21 21 0.000e+00 PASS
MOT17-04-FRCNN partially_tracked 36 36 0.000e+00 PASS
MOT17-04-FRCNN mostly_lost 12 12 0.000e+00 PASS
MOT17-04-FRCNN mtr 0.304347826087 0.304347826087 0.000e+00 PASS
MOT17-04-FRCNN ptr 0.521739130435 0.521739130435 0.000e+00 PASS
MOT17-04-FRCNN mlr 0.173913043478 0.173913043478 0.000e+00 PASS
MOT17-04-FRCNN num_false_positives 269 269 0.000e+00 PASS
MOT17-04-FRCNN num_misses 9650 9650 0.000e+00 PASS
MOT17-04-FRCNN num_switches 482 482 0.000e+00 PASS
MOT17-04-FRCNN num_fragmentations 910 910 0.000e+00 PASS
MOT17-04-FRCNN mota 0.569815534784 0.569815534784 0.000e+00 PASS
MOT17-04-FRCNN moda 0.589751013318 0.589751013318 0.000e+00 PASS
MOT17-04-FRCNN motp 0.127817758341 0.127817758341 8.327e-17 PASS
MOT17-04-FRCNN smota 0.493012805311 0.493012805311 1.110e-16 PASS
MOT17-04-FRCNN clr_f1 0.745503527903 0.745503527903 0.000e+00 PASS
MOT17-04-FRCNN fp_per_frame 0.513358778626 0.513358778626 0.000e+00 PASS
MOT17-04-FRCNN hota 0.470316535354 0.470316535354 2.776e-16 PASS
MOT17-04-FRCNN deta 0.519342274572 0.519342274572 0.000e+00 PASS
MOT17-04-FRCNN assa 0.431613360756 0.431613360756 5.551e-16 PASS
MOT17-04-FRCNN detre 0.537759424618 0.537759424618 0.000e+00 PASS
MOT17-04-FRCNN detpr 0.878688069772 0.878688069772 0.000e+00 PASS
MOT17-04-FRCNN assre 0.440294164372 0.440294164372 4.996e-16 PASS
MOT17-04-FRCNN asspr 0.924313727813 0.924313727813 2.220e-16 PASS
MOT17-04-FRCNN loca 0.886341100755 0.886341100755 2.220e-16 PASS
MOT17-04-FRCNN owta 0.479927255879 0.479927255879 2.776e-16 PASS
MOT17-05-FRCNN idf1 0.0896696381751 0.0896696381751 0.000e+00 PASS
MOT17-05-FRCNN idp 0.374179431072 0.374179431072 0.000e+00 PASS
MOT17-05-FRCNN idr 0.0509383378016 0.0509383378016 0.000e+00 PASS
MOT17-05-FRCNN recall 0.13345248734 0.13345248734 0.000e+00 PASS
MOT17-05-FRCNN precision 0.980306345733 0.980306345733 0.000e+00 PASS
MOT17-05-FRCNN num_unique_objects 71 71 0.000e+00 PASS
MOT17-05-FRCNN mostly_tracked 0 0 0.000e+00 PASS
MOT17-05-FRCNN partially_tracked 11 11 0.000e+00 PASS
MOT17-05-FRCNN mostly_lost 60 60 0.000e+00 PASS
MOT17-05-FRCNN mtr 0 0 0.000e+00 PASS
MOT17-05-FRCNN ptr 0.154929577465 0.154929577465 0.000e+00 PASS
MOT17-05-FRCNN mlr 0.845070422535 0.845070422535 0.000e+00 PASS
MOT17-05-FRCNN num_false_positives 9 9 0.000e+00 PASS
MOT17-05-FRCNN num_misses 2909 2909 0.000e+00 PASS
MOT17-05-FRCNN num_switches 106 106 0.000e+00 PASS
MOT17-05-FRCNN num_fragmentations 109 109 0.000e+00 PASS
MOT17-05-FRCNN mota 0.0991957104558 0.0991957104558 4.163e-17 PASS
MOT17-05-FRCNN moda 0.130771522192 0.130771522192 0.000e+00 PASS
MOT17-05-FRCNN motp 0.140782733337 0.140782733337 2.776e-17 PASS
MOT17-05-FRCNN smota 0.0804079045174 0.0804079045174 0.000e+00 PASS
MOT17-05-FRCNN clr_f1 0.234923964342 0.234923964342 0.000e+00 PASS
MOT17-05-FRCNN fp_per_frame 0.0215311004785 0.0215311004785 0.000e+00 PASS
MOT17-05-FRCNN hota 0.0954591007631 0.0954591007631 4.163e-17 PASS
MOT17-05-FRCNN deta 0.116690804588 0.116690804588 0.000e+00 PASS
MOT17-05-FRCNN assa 0.0782294415609 0.0782294415609 5.551e-17 PASS
MOT17-05-FRCNN detre 0.117633225154 0.117633225154 0.000e+00 PASS
MOT17-05-FRCNN detpr 0.864102268801 0.864102268801 0.000e+00 PASS
MOT17-05-FRCNN assre 0.0794302241527 0.0794302241527 5.551e-17 PASS
MOT17-05-FRCNN asspr 0.928636949524 0.928636949524 0.000e+00 PASS
MOT17-05-FRCNN loca 0.871531959574 0.871531959574 0.000e+00 PASS
MOT17-05-FRCNN owta 0.0958406283428 0.0958406283428 4.163e-17 PASS
MOT17-09-FRCNN idf1 0.0958046336882 0.0958046336882 0.000e+00 PASS
MOT17-09-FRCNN idp 0.485714285714 0.485714285714 0.000e+00 PASS
MOT17-09-FRCNN idr 0.0531434525877 0.0531434525877 0.000e+00 PASS
MOT17-09-FRCNN recall 0.109065647794 0.109065647794 0.000e+00 PASS
MOT17-09-FRCNN precision 0.996825396825 0.996825396825 0.000e+00 PASS
MOT17-09-FRCNN num_unique_objects 22 22 0.000e+00 PASS
MOT17-09-FRCNN mostly_tracked 0 0 0.000e+00 PASS
MOT17-09-FRCNN partially_tracked 2 2 0.000e+00 PASS
MOT17-09-FRCNN mostly_lost 20 20 0.000e+00 PASS
MOT17-09-FRCNN mtr 0 0 0.000e+00 PASS
MOT17-09-FRCNN ptr 0.0909090909091 0.0909090909091 0.000e+00 PASS
MOT17-09-FRCNN mlr 0.909090909091 0.909090909091 0.000e+00 PASS
MOT17-09-FRCNN num_false_positives 1 1 0.000e+00 PASS
MOT17-09-FRCNN num_misses 2565 2565 0.000e+00 PASS
MOT17-09-FRCNN num_switches 55 55 0.000e+00 PASS
MOT17-09-FRCNN num_fragmentations 61 61 0.000e+00 PASS
MOT17-09-FRCNN mota 0.0896144494616 0.0896144494616 5.551e-17 PASS
MOT17-09-FRCNN moda 0.108718304967 0.108718304967 0.000e+00 PASS
MOT17-09-FRCNN motp 0.162019146505 0.162019146505 1.110e-16 PASS
MOT17-09-FRCNN smota 0.071943726293 0.071943726293 0.000e+00 PASS
MOT17-09-FRCNN clr_f1 0.196618659987 0.196618659987 0.000e+00 PASS
MOT17-09-FRCNN fp_per_frame 0.00381679389313 0.00381679389313 0.000e+00 PASS
MOT17-09-FRCNN hota 0.0717384178417 0.0717384178417 1.388e-17 PASS
MOT17-09-FRCNN deta 0.0922693243415 0.0922693243415 0.000e+00 PASS
MOT17-09-FRCNN assa 0.0566116465384 0.0566116465384 2.082e-17 PASS
MOT17-09-FRCNN detre 0.0928319409152 0.0928319409152 0.000e+00 PASS
MOT17-09-FRCNN detpr 0.848454469507 0.848454469507 0.000e+00 PASS
MOT17-09-FRCNN assre 0.0569184672945 0.0569184672945 2.082e-17 PASS
MOT17-09-FRCNN asspr 0.92190862671 0.92190862671 0.000e+00 PASS
MOT17-09-FRCNN loca 0.854155996663 0.854155996663 0.000e+00 PASS
MOT17-09-FRCNN owta 0.0720174507895 0.0720174507895 1.388e-17 PASS
MOT17-10-FRCNN idf1 0.0706245181187 0.0706245181187 0.000e+00 PASS
MOT17-10-FRCNN idp 0.407473309609 0.407473309609 0.000e+00 PASS
MOT17-10-FRCNN idr 0.0386628397771 0.0386628397771 0.000e+00 PASS
MOT17-10-FRCNN recall 0.0945466824244 0.0945466824244 0.000e+00 PASS
MOT17-10-FRCNN precision 0.996441281139 0.996441281139 0.000e+00 PASS
MOT17-10-FRCNN num_unique_objects 36 36 0.000e+00 PASS
MOT17-10-FRCNN mostly_tracked 0 0 0.000e+00 PASS
MOT17-10-FRCNN partially_tracked 5 5 0.000e+00 PASS
MOT17-10-FRCNN mostly_lost 31 31 0.000e+00 PASS
MOT17-10-FRCNN mtr 0 0 0.000e+00 PASS
MOT17-10-FRCNN ptr 0.138888888889 0.138888888889 0.000e+00 PASS
MOT17-10-FRCNN mlr 0.861111111111 0.861111111111 0.000e+00 PASS
MOT17-10-FRCNN num_false_positives 2 2 0.000e+00 PASS
MOT17-10-FRCNN num_misses 5363 5363 0.000e+00 PASS
MOT17-10-FRCNN num_switches 104 104 0.000e+00 PASS
MOT17-10-FRCNN num_fragmentations 134 134 0.000e+00 PASS
MOT17-10-FRCNN mota 0.0766503461084 0.0766503461084 4.163e-17 PASS
MOT17-10-FRCNN moda 0.0942090157015 0.0942090157015 0.000e+00 PASS
MOT17-10-FRCNN motp 0.189341848837 0.189341848837 3.331e-16 PASS
MOT17-10-FRCNN smota 0.0587487024567 0.0587487024567 2.776e-17 PASS
MOT17-10-FRCNN clr_f1 0.172706245181 0.172706245181 0.000e+00 PASS
MOT17-10-FRCNN fp_per_frame 0.00613496932515 0.00613496932515 0.000e+00 PASS
MOT17-10-FRCNN hota 0.0874803453103 0.0874803453103 0.000e+00 PASS
MOT17-10-FRCNN deta 0.0780272459924 0.0780272459924 0.000e+00 PASS
MOT17-10-FRCNN assa 0.100391264243 0.100391264243 0.000e+00 PASS
MOT17-10-FRCNN detre 0.0784364253534 0.0784364253534 0.000e+00 PASS
MOT17-10-FRCNN detpr 0.826652931261 0.826652931261 0.000e+00 PASS
MOT17-10-FRCNN assre 0.101174910367 0.101174910367 0.000e+00 PASS
MOT17-10-FRCNN asspr 0.889110587253 0.889110587253 0.000e+00 PASS
MOT17-10-FRCNN loca 0.836230808969 0.836230808969 0.000e+00 PASS
MOT17-10-FRCNN owta 0.0877811382222 0.0877811382222 0.000e+00 PASS
MOT17-11-FRCNN idf1 0.275459098497 0.275459098497 0.000e+00 PASS
MOT17-11-FRCNN idp 0.560081466395 0.560081466395 0.000e+00 PASS
MOT17-11-FRCNN idr 0.182643347354 0.182643347354 0.000e+00 PASS
MOT17-11-FRCNN recall 0.312818242196 0.312818242196 0.000e+00 PASS
MOT17-11-FRCNN precision 0.959266802444 0.959266802444 0.000e+00 PASS
MOT17-11-FRCNN num_unique_objects 44 44 0.000e+00 PASS
MOT17-11-FRCNN mostly_tracked 2 2 0.000e+00 PASS
MOT17-11-FRCNN partially_tracked 17 17 0.000e+00 PASS
MOT17-11-FRCNN mostly_lost 25 25 0.000e+00 PASS
MOT17-11-FRCNN mtr 0.0454545454545 0.0454545454545 0.000e+00 PASS
MOT17-11-FRCNN ptr 0.386363636364 0.386363636364 0.000e+00 PASS
MOT17-11-FRCNN mlr 0.568181818182 0.568181818182 0.000e+00 PASS
MOT17-11-FRCNN num_false_positives 60 60 0.000e+00 PASS
MOT17-11-FRCNN num_misses 3104 3104 0.000e+00 PASS
MOT17-11-FRCNN num_switches 73 73 0.000e+00 PASS
MOT17-11-FRCNN num_fragmentations 137 137 0.000e+00 PASS
MOT17-11-FRCNN mota 0.283373920744 0.283373920744 5.551e-17 PASS
MOT17-11-FRCNN moda 0.299535089661 0.299535089661 0.000e+00 PASS
MOT17-11-FRCNN motp 0.113982215708 0.113982215708 6.939e-16 PASS
MOT17-11-FRCNN smota 0.247718204385 0.247718204385 1.943e-16 PASS
MOT17-11-FRCNN clr_f1 0.471786310518 0.471786310518 0.000e+00 PASS
MOT17-11-FRCNN fp_per_frame 0.133630289532 0.133630289532 0.000e+00 PASS
MOT17-11-FRCNN hota 0.257713560364 0.257713560364 5.551e-17 PASS
MOT17-11-FRCNN deta 0.279233950654 0.279233950654 0.000e+00 PASS
MOT17-11-FRCNN assa 0.238577199402 0.238577199402 1.943e-16 PASS
MOT17-11-FRCNN detre 0.285797513487 0.285797513487 0.000e+00 PASS
MOT17-11-FRCNN detpr 0.876406903205 0.876406903205 0.000e+00 PASS
MOT17-11-FRCNN assre 0.243341703415 0.243341703415 1.665e-16 PASS
MOT17-11-FRCNN asspr 0.92575082276 0.92575082276 0.000e+00 PASS
MOT17-11-FRCNN loca 0.891956208658 0.891956208658 7.772e-16 PASS
MOT17-11-FRCNN owta 0.260926488915 0.260926488915 1.110e-16 PASS
MOT17-13-FRCNN idf1 0.0711935387377 0.0711935387377 0.000e+00 PASS
MOT17-13-FRCNN idp 0.636363636364 0.636363636364 0.000e+00 PASS
MOT17-13-FRCNN idr 0.0377059569075 0.0377059569075 0.000e+00 PASS
MOT17-13-FRCNN recall 0.058618504436 0.058618504436 0.000e+00 PASS
MOT17-13-FRCNN precision 0.989304812834 0.989304812834 0.000e+00 PASS
MOT17-13-FRCNN num_unique_objects 44 44 0.000e+00 PASS
MOT17-13-FRCNN mostly_tracked 0 0 0.000e+00 PASS
MOT17-13-FRCNN partially_tracked 4 4 0.000e+00 PASS
MOT17-13-FRCNN mostly_lost 40 40 0.000e+00 PASS
MOT17-13-FRCNN mtr 0 0 0.000e+00 PASS
MOT17-13-FRCNN ptr 0.0909090909091 0.0909090909091 0.000e+00 PASS
MOT17-13-FRCNN mlr 0.909090909091 0.909090909091 0.000e+00 PASS
MOT17-13-FRCNN num_false_positives 2 2 0.000e+00 PASS
MOT17-13-FRCNN num_misses 2971 2971 0.000e+00 PASS
MOT17-13-FRCNN num_switches 41 41 0.000e+00 PASS
MOT17-13-FRCNN num_fragmentations 47 47 0.000e+00 PASS
MOT17-13-FRCNN mota 0.0449936628644 0.0449936628644 3.469e-17 PASS
MOT17-13-FRCNN moda 0.0579847908745 0.0579847908745 0.000e+00 PASS
MOT17-13-FRCNN motp 0.193720373223 0.193720373223 1.665e-16 PASS
MOT17-13-FRCNN smota 0.0336380643073 0.0336380643073 1.388e-17 PASS
MOT17-13-FRCNN clr_f1 0.110679030811 0.110679030811 0.000e+00 PASS
MOT17-13-FRCNN fp_per_frame 0.00534759358289 0.00534759358289 0.000e+00 PASS
MOT17-13-FRCNN hota 0.0790953463109 0.0790953463109 0.000e+00 PASS
MOT17-13-FRCNN deta 0.0482024139453 0.0482024139453 0.000e+00 PASS
MOT17-13-FRCNN assa 0.134592138082 0.134592138082 0.000e+00 PASS
MOT17-13-FRCNN detre 0.0484123807618 0.0484123807618 0.000e+00 PASS
MOT17-13-FRCNN detpr 0.817056009006 0.817056009006 0.000e+00 PASS
MOT17-13-FRCNN assre 0.136125737666 0.136125737666 0.000e+00 PASS
MOT17-13-FRCNN asspr 0.918090220879 0.918090220879 0.000e+00 PASS
MOT17-13-FRCNN loca 0.834113987906 0.834113987906 1.110e-16 PASS
MOT17-13-FRCNN owta 0.0793300823194 0.0793300823194 0.000e+00 PASS
OVERALL idf1 0.330490492683 0.330490492683 0.000e+00 PASS
OVERALL idp 0.630043321676 0.630043321676 0.000e+00 PASS
OVERALL idr 0.223993319725 0.223993319725 0.000e+00 PASS
OVERALL recall 0.348190758953 0.348190758953 0.000e+00 PASS
OVERALL precision 0.979383057571 0.979383057571 0.000e+00 PASS
OVERALL num_unique_objects 339 339 0.000e+00 PASS
OVERALL mostly_tracked 23 23 0.000e+00 PASS
OVERALL partially_tracked 83 83 0.000e+00 PASS
OVERALL mostly_lost 233 233 0.000e+00 PASS
OVERALL mtr 0.0678466076696 0.0678466076696 0.000e+00 PASS
OVERALL ptr 0.244837758112 0.244837758112 0.000e+00 PASS
OVERALL mlr 0.687315634218 0.687315634218 0.000e+00 PASS
OVERALL num_false_positives 395 395 0.000e+00 PASS
OVERALL num_misses 35126 35126 0.000e+00 PASS
OVERALL num_switches 1100 1100 0.000e+00 PASS
OVERALL num_fragmentations 1674 1674 0.000e+00 PASS
OVERALL mota 0.320449062906 0.320449062906 0.000e+00 PASS
OVERALL moda 0.340861013175 0.340861013175 0.000e+00 PASS
OVERALL motp 0.131261153213 0.131261153213 8.327e-17 PASS
OVERALL smota 0.274745142347 0.274745142347 5.551e-17 PASS
OVERALL clr_f1 0.513737354379 0.513737354379 0.000e+00 PASS
OVERALL fp_per_frame 0.148944193062 0.148944193062 0.000e+00 PASS
OVERALL hota 0.333427606588 0.333427606588 1.665e-16 PASS
OVERALL deta 0.303700764938 0.303700764938 0.000e+00 PASS
OVERALL assa 0.370854628363 0.370854628363 3.886e-16 PASS
OVERALL detre 0.310530222383 0.310530222383 0.000e+00 PASS
OVERALL detpr 0.873452355771 0.873452355771 0.000e+00 PASS
OVERALL assre 0.378413164882 0.378413164882 4.996e-16 PASS
OVERALL asspr 0.924557170786 0.924557170786 2.220e-16 PASS
OVERALL loca 0.882723874929 0.882723874929 2.220e-16 PASS
OVERALL owta 0.337801742209 0.337801742209 2.776e-16 PASS

@mikel-brostrom
mikel-brostrom marked this pull request as ready for review July 19, 2026 23:20
@cheind

cheind commented Aug 1, 2026

Copy link
Copy Markdown
Owner

@mikel-brostrom I had an in depth-look and I lilke the simplification. I will accept the PR once I'm back from vacation! Thanks!

@mikel-brostrom

Copy link
Copy Markdown
Contributor Author

Sorry for my late response @cheind. I have also been on vacation 😅. Glad to see that this repo will get a proper update.

@mikel-brostrom

mikel-brostrom commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor Author

I added some important information to the README for how to reproduce trackeval evaluation results, which take into consideration distractor classes among other things

@mikel-brostrom
mikel-brostrom marked this pull request as draft August 31, 2026 09:58
@mikel-brostrom
mikel-brostrom marked this pull request as ready for review August 31, 2026 11:29
@0xamlab

0xamlab commented Sep 20, 2026

Copy link
Copy Markdown

This PR is very close to something we've been building, so sharing measurements in case they're useful. Disclosure up front: I work on this at thyn-ai.

We built motmetrics-mojo — a clean-room Mojo kernel that computes the MOTChallenge metric set (MOTA, MOTP, IDF1, IDP, IDR, MOSTLY_TRACKED / PARTIALLY_TRACKED / MOSTLY_LOST, precision, recall, fragmentation, and the supporting counts) as a single-pass event-stream accumulator, without materializing the per-frame pandas event table. It mirrors the released motmetrics API (MOTAccumulator + compute), with motmetrics 1.4.0 (default scipy assignment-solver path) as the correctness oracle.

Correctness evidence (differential suite: seeded synthetic MOTChallenge-style streams plus adversarial cases — tie-heavy integer distance matrices, all-NaN matrices, empty and one-sided frames, fragmentation boundaries, max_switch_time variants):

  • exact equality on every count and every integer-derived metric (MOTA, IDF1, IDP, IDR, precision, recall)
  • MOTP agrees to a measured 6.2e-15 (its distance sum is the only floating-point accumulation)
  • both the native backend and the vendored pure-Python fallback are tested against the oracle, and the two backends agree with each other bit-for-bit

Benchmarks (macOS arm64, Apple M4 Max; Python 3.12.14, NumPy 2.5.3, motmetrics 1.4.0; a pipeline is accumulate a whole seeded stream frame by frame, then compute the 18 MOTChallenge metrics; correctness against the oracle is asserted before timing):

pipeline py-motmetrics 1.4.0 motmetrics-mojo speedup
cold, 1,000 frames x 40 objects x 44 hypotheses, fresh process 2.055 s 0.0551 s 37.3x
warm, 100 frames 0.033 s 0.00086 s 38.4x
warm, 500 frames 0.456 s 0.00997 s 45.7x
warm, 1,000 frames 2.000 s 0.04839 s 41.3x
warm, 2,000 frames 20.473 s 0.42267 s 48.4x

i.e. 37–48x on 100–2,000-frame streams, cold and warm alike.

Package shape: per-platform wheels (macOS arm64, Linux x86_64) carry the compiled kernel with the Mojo runtime vendored in, so no toolchain is needed; on every other platform the same API transparently runs on the vendored pure-Python fallback. NumPy is the only runtime dependency.

Since this PR rebuilds evaluation around one fast path: if there's interest, we'd be glad to collaborate on making the kernel an official optional fast path behind the motmetrics API (e.g. an extra such as motmetrics[mojo]), with motmetrics remaining the reference implementation and the fallback for everything outside the kernel's scope. Benchmark script, differential suite, and wheel build are in the repo: https://github.com/thyn-ai/mojo-kernels/tree/chore/kernel-motmetrics/python/motmetrics_mojo

@0xamlab

0xamlab commented Sep 20, 2026

Copy link
Copy Markdown

Update: the package has landed on main — https://github.com/thyn-ai/mojo-kernels/tree/main/python/motmetrics_mojo

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants