My object detector was too slow for live video. So I stopped waiting for it.

A camera at 30 frames a second gives you a new frame every 33 ms. The object detector I wanted to use takes about 80 ms on my laptop CPU. So by the time it finishes one frame, the camera has moved on by two more.

Most tracking code ignores this. It runs the detector on every frame and assumes the detector keeps up. On recorded video that is fine, because the video waits. On a live stream, the world does not wait, and the boxes you draw are for where people were, not where they are.

I built leantrack to answer one question: if the detector is a limited resource, when should it run, and what does each choice cost in accuracy?

Two views of the same street. On the left the boxes trail behind the people. On the right they stay on them.

On the left, the loop waits for the detector, so the boxes are behind the people. On the right, the detector runs in the background and the boxes stay on them. Same detector, same video.

The setup

A tracker does two jobs. The detector finds people in one frame. The tracker connects those detections across frames, so that person 7 stays person 7.

Two cautions before any numbers. First, my detector has general-purpose weights and was never trained on this dataset, so the absolute scores are low. The comparisons between configurations are the result. Second, I wrote the tracker core myself, so I checked it against a known implementation first: on the same detections it scores 47.11 HOTA against 47.32 for BoxMOT's ByteTrack. Close enough that the experiments after it mean something.

First idea: run the detector less often

The obvious move is to run the detector on every Nth frame and let the tracker predict the frames in between.

Detector runs on Mean time per frame HOTA
Every frame 79.8 ms 37.5
Every 3rd frame 26.7 ms 36.0
Every 5th frame 16.1 ms 33.0
Every 10th frame 8.1 ms 28.9

Skipping two frames out of three cuts the mean time by 67% and costs 1.5 points. That is a good trade. Past that, accuracy falls quickly.

There is a catch in this table that the mean hides. The slowest frames did not get faster: the 99th percentile stayed near 80 ms at every interval. A frame that runs the detector costs what it always cost. A fixed interval improves your average and does nothing for your worst case.

Second idea: watch the pixels between detector runs

Between detector runs, the tracker is guessing. It assumes each person keeps moving the way they were. I added optical flow, which follows small patches of the image from one frame to the next, and gave that motion to the tracker as a measurement.

HOTA against mean frame time, with and without optical flow. The line with optical flow is higher, and the gap grows at long intervals.

It costs less than 2 ms a frame. The gain grows with the interval: 0.3 points at every 3rd frame, 3.2 points at every 10th, 6.7 points at every 20th. It also cut identity switches by a third to a half. With flow, detecting every 10th frame is almost as good as detecting every 5th frame without it.

The real problem: a live stream

Everything above processes every frame and ignores the clock. So I simulated one. Frames arrive at the camera's rate, and an output only counts if it is ready before the next frame arrives.

I compared three loops:

Loop Mean output latency HOTA
Blocking 101.4 ms 32.9
Background 5.1 ms 34.5
Background, no correction 2.8 ms 31.6

The background loop answers about 20 times sooner (101.4 / 5.1 = 19.9) and scores 1.6 points higher. It is faster and better, which does not happen often.

The third row is the part I would have got wrong without measuring. A detection that arrives three frames late describes where people were three frames ago. Use it as it is and the score drops 2.9 points, below even the blocking loop. Moving the late boxes forward with optical flow is what makes the background design work.

Two results I did not expect

The biggest model was not the best one. Offline, the larger YOLOX-m beats YOLOX-s, 40.0 to 37.5. On the live stream, in the background loop, it loses: 33.0 to 34.5. Its results arrive five frames old instead of three, and a better answer about the distant past is worth less than a decent answer about the recent past.

A faster model beat a smarter schedule. I quantized YOLOX-s to 8-bit integers. That made it 2.8 times faster and cost 1.0 point offline. On the live stream it now fit inside one frame period, and a plain blocking loop with the quantized model scored 36.4, which is 1.9 points better than my careful background loop with the original model. All the scheduling work was worth less than making the detector fast enough to not need it.

What did not work

These are in the repository, because they changed the design.

Triggering the detector when the tracker looks unsure. This seemed like the smart version of a fixed interval. I built two uncertainty signals, and they really do predict a bad box: when flow reliability was low, the box was wrong 78% of the time, against 16% when it was high. But running the detector at those moments did not fix the box. The trigger landed within half a point of a fixed interval, or below it. My guess is that low reliability often means the person is hidden, and a detector cannot see a hidden person either. I did not test that.

Predicting which boxes are wrong. A small model on eight features reached an AUC of 0.73. One of those features alone reached 0.71. And its probabilities were off on scenes with a moving camera, so filtering on them made the score worse.

The background loop for a fast detector. If the detector already fits in the frame period, moving it to the background just makes its answer one frame late. A blocking loop was 0.6 points better. I added a frame budget for this: wait for the detector only if it is expected to finish in time.

One more thing: getting a lost person back

When someone walks behind a pole, the tracker loses them. When they come out, do they get their old identity back?

Share of tracks that recover against the length of the gap, for four methods. The two appearance methods stay far above the two without appearance.

With position alone, 16% of tracks recovered after a 30-frame gap. Comparing appearance as well raised that to 57%. The surprise was how little was needed: a plain color histogram matched a neural re-identification network for gaps up to 30 frames, at a small fraction of the cost (0.07 ms a box against 1.6 ms). The network only pulled ahead on longer gaps.

And then the honest part: on the standard benchmark with strong detections, this changed the overall score by less than 0.1. The capability is real and the benchmark does not reward it.

What I took from this

Limits

One laptop CPU, one run per configuration, no GPU and no small device. The live stream experiments use a simulated clock with measured timings, not a real camera. Changing the detector timing alone moves HOTA by about 0.7 on a live stream, so smaller differences between two live configurations are noise.

The code, every table, and eight short design records are in the repository.

All posts