Nous poursuivons notre apprentissage de votre langue

Nous travaillons dur pour que toutes les pages de milestonesys.com soient disponibles dans autant de langues que possible. Mais c’est un processus qui requiert du temps. En attendant, un grand nombre de nos fonctions sont déjà proposées en plusieurs langues. Certaines pages, comme celle-ci, ne sont pas encore disponibles dans votre langue.
Merci de votre compréhension.

What happens when traffic AI model meets a new city?

août 14, 2026

A traffic detection model that performs well in one city may be 50% less accurate in another.

That was one of the clearest findings from Track 6 of the 10th AI City Challenge. Teams trained object detection models on traffic video from a U.S. city and were evaluated on a hidden benchmark that included a European city, a scenario they were not trained for.  

When evaluated against the unknown city, overall detection accuracy dropped 43%. For smaller objects like distant pedestrians, cyclists and far-away vehicles, the drop was 55%. 

For teams building computer vision systems for real-world situations, that finding matters: a model that performs well in one environment likely behaves differently in another. 

The result highlights a challenge many AI teams face – the need for access to real-world video data to build models that perform reliably in the specific environments and scenarios where accuracy is essential.

Why data from one city is not enough

Computer vision models learn visual patterns present in their training data. In traffic video training data, that includes camera placement, road design, vehicle mix, traffic density, lighting, weather, and object scale. When you change those conditions and model performance can change with them. 

Consider a pedestrian crossing a wide, open intersection who is recorded by a camera mounted high on a pole. In this view, the pedestrian occupies a large, unobstructed part of the frame. Now consider the same pedestrian on a narrower street, recorded from a lower position with parked vehicles in the camera’s line of sight. In this view, the pedestrian is smaller, partly hidden, and seen from a different angle. "Pedestrian” is still the same object class, but the visual context has changed significantly. 

This is why a model that performs well on the data it was trained on won’t necessarily perform as well once conditions change. It is particularly relevant for traffic and mobility, where conditions are never identical from one city to the next. 

But accessing the right data for training specialized models is often difficult.

Hafnia provides the infrastructure and services to make relevant real-world video data usable for AI development. It helps teams prepare data for model training and gives developers a managed way to train and evaluate models against footage that reflects the environments they are building for.  

The test: train in one city, perform across two

Track 6 focused on fine-grained object detection when the deployment environment differs from the training environment. The challenge was to build a model that performs well beyond the environment represented in the available training data.  

Track 6 was organized by Milestone through Hafnia, together with NVIDIA and Universidad Autónoma de Madrid. The track received 127 registrations, with 100 participants approved and 86 becoming active during the competition. Participants represented 23 countries, 75 research institutions and 17 companies, and ran 1,926 training experiments during the challenge. 

Participants trained on approximately 13,000 frames containing around 150,000 annotated objects. The task included 10 object classes covering people, vehicles, and other road users. 

The hidden benchmark included data from both the US city, where the training data came from, and a second European city with different visual and geographic characteristics. The final rankings were based on performance across the full benchmark. 

The result: geography changed the model’s performance

In a reference evaluation, overall detection performance dropped 43% when models moved across geographically different cities. Small-object detection dropped 55%. 

Metric  Source city  Second city  Change 
mAP  0.517  0.294  -43% 
mAP50  0.685  0.398  -42% 
Small objects  0.094  0.042  -55% 
Large objects  0.639  0.404  -37% 

 

The performance gap was visible across all metrics. Small-object detection experienced the largest decline, while large-object detection also fell substantially. 

What this shows isn't a quirk of two specific cities. It's that the model never learned to handle conditions outside its training data.

Participants reported the same pattern. Several teams found that improvements on the available training data did not translate into better results on the hidden benchmark.

The biggest technical challenge for us was the performance degradation that occurred when our models were transferred from the training domain to the hidden target domain. It was difficult to determine how well a model would generalize to unseen data, and improvements on the available training or validation data did not always lead to better performance on the hidden domain.
School of Optics and Photonics, Beijing Institute of Technology

AI teams building production systems should take note: performance on a single dataset can hide weaknesses that only appear when geography, infrastructure, and camera conditions change. 

How Training-as-a-service supported the challenge

Hafnia Training-as-a-Service provided the managed environment participants in Track 6 used to train against the challenge data. 

Instead of downloading the complete dataset, participants packaged their models, training code, configurations, and dependencies into a trainer that ran inside the Hafnia environment. 

This allowed external teams to fine-tune their models on real-world video data without receiving direct access to the underlying dataset. For participants, that meant they could train on data they would otherwise not be able to access on their own, without having to source, transfer, or manage it themselves. 

Hafnia offers a growing library of real-world video data, plus the infrastructure to curate, anonymize, train and evaluate against it. 

What comes next

The AI City Challenge continues at ECCV 2026 in Malmö, where participating teams will present their work and share lessons from Track 6. Milestone Hafnia will also be there to share more from the challenge and the work underway across the platform.

The results offer a useful reminder for computer vision developers as they move from development to deployment: models perform best when they are trained and evaluated on data that reflects the conditions where they are expected to operate. 

Are you building computer vision models for real world deployment? 

Talk to the Hafnia team about Training-as-a-Service and early access the platform. 

Tags
Ready to see what we have to offer with smart video technology?
Book a demo
You will be logged out in
5 minutes and 0 seconds
For your security, sessions automatically end after 15 minutes of inactivity unless you choose to stay logged in.