Delineate Anything v2: A Global Foundation Model for Field Delineation
AuthorsMykola Lavreniuk, Nataliia Kussul, Andrii Shelestov, Yevhenii Salii, Volodymyr Kuzin, Charlotte Julia Li-Xing Wang, Zoltan Szantoi
A globally scalable computer vision foundation model maps agricultural field boundaries across countries with major accuracy gains and practical nationwide deployment speed.
Key results
Multi-resolution field-boundary dataset spanning 61 countries.
Independent evaluation benchmark covering 100 countries.
Benchmark score for the YOLOv11-based model.
Improvement over the original Delineate Anything framework.
Hours required to map 603K km² on a consumer-grade workstation.
What the paper found
Researchers involving the European Space Agency introduce Delineate Anything v2, a global foundation model for agricultural field-boundary delineation that targets the main failure of satellite segmentation: structural label noise caused when administrative parcels merge multiple physical fields. Rather than redesigning the network, the authors build FBIS-73M, a multi-resolution dataset containing 73M field instances across 61 countries, and curate it with resolution-specific remediation. High-resolution anomalies receive manual geometric splitting, while medium-resolution imagery uses HSV pixel homogenization to suppress false internal boundaries and Khalimsky-grid edge enhancement to strengthen genuine but visually weak borders. A YOLOv11-based instance-segmentation model is evaluated on an independent benchmark spanning 100 countries, where Delineate Anything v2 reaches 0.559 mAP@0.5, compared with 0.275 for the original Delineate Anything, a 103.3% relative improvement. Ablations show that curation, especially medium-resolution image-space remediation, contributes more than simply scaling raw data. The system also demonstrates operational scalability by mapping Ukraine’s 603K km² in 5.4 hours on a consumer-grade workstation. The work contrasts with general vision foundation models such as Meta’s Segment Anything Model, which often lack the scale and topological awareness required for geospatial field boundaries, and argues that supervision quality and geographic diversity are more important bottlenecks than model complexity.
Original abstract
Accurate agricultural field boundary delineation at large scale is a foundational task for food security, supply chain transparency, and carbon accounting. While vision foundation models like SAM show remarkable zero-shot capabilities, they frequently fail in geospatial domains due to topological complexity, cropland texturing patterns, and a lack of physical scale awareness. In this work, we introduce Delineate Anything v2, a globally scalable foundation model designed specifically for wide-area field boundary mapping. We construct FBIS-73M, a 73-million-instance multi-resolution dataset spanning 61 countries. To address the pervasive issue of multi-field administrative parcel merging, we introduce a resolution-specific data curation pipeline that leverages topological image-space adaptation to homogenize merged parcels and strengthen weak physical boundaries. Furthermore, we establish a novel, manually curated evaluation benchmark covering 100 countries to assess independent zero-shot generalization. Our results show that Delineate Anything v2 surpasses the current state-of-the-art, including the Delineate Anything framework, by 0.284 mAP@0.5 (+103.3% relative gain), while maintaining execution speeds suitable for rapid national- and global-scale deployment, as demonstrated by nationwide mapping of Ukraine (603,000 km^2) in 5.4 hours on a consumer-grade workstation. Code, pre-trained weights, the FBIS-73M dataset, and ready-to-use national-scale vector boundary products are publicly available at https://github.com/Lavreniuk/Delineate-Anything.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.