OVEarth-Bench is a Python evaluation suite for remote-sensing image understanding. It supports three tasks: open-vocabulary segmentation/detection (open_vocabulary), referring-expression segmentation/grounding (referring), and reasoning segmentation/grounding (reasoning).
- Installation
- Project Layout
- Tasks
- Data Formats
- Command-Line Interface
- Python API
- Metrics
- Output JSON
- Negative Queries
- Oriented Bounding Box Evaluation
- Tests
- Dependencies
git clone <repo-url>
cd OVEarth-bench
# Standard installation
pip install .
# Editable installation with test dependencies
pip install -e ".[dev]"After installation, the ovearth-eval command is available:
ovearth-eval --help(project root)/
|-- pyproject.toml
|-- README.md
|-- README-ZH.md
|-- ovearth_eval/
| |-- __init__.py
| |-- _utils.py
| |-- types.py
| |-- presence.py
| |-- segmentation.py
| |-- detection.py
| |-- oriented_detection.py
| |-- polygon.py
| |-- rle.py
| |-- open_vocab.py
| |-- io.py
| `-- cli.py
`-- tests/
|-- test_presence.py
|-- test_segmentation.py
|-- test_detection.py
|-- test_polygon.py
|-- test_rle.py
|-- test_open_vocab.py
|-- test_io.py
`-- test_cli.py
Select a task with --task:
| Task | CLI value | Prediction list | Samples | Task-specific metrics |
|---|---|---|---|---|
| Open-vocabulary segmentation/detection | vocabulary |
tasks.open_vocabulary |
Positive and negative | Presence metrics and negative-query false-positive rate |
| Referring-expression segmentation/localization | referring |
tasks.referring |
Positive only | None |
| Reasoning segmentation/localization | reasoning |
tasks.reasoning |
Positive only | None |
- Open vocabulary: Given a category phrase such as
"bell tower", the model must determine whether that category is present and then segment or localize it. Negative queries measure hallucination suppression. - Referring: Given a spatial or contextual description, the model must segment or localize the referred object.
- Reasoning: Given a functional or attribute-based description, the model must infer the target and segment or localize it.
The current ground_truth.json contains 590 annotations, 172 categories, and 5,023 queries:
| Task | Queries | Positive | Negative | query_id range |
|---|---|---|---|---|
open_vocabulary |
3,235 | 1,067 | 2,168 | ov_000001 - ov_003235 |
referring |
732 | 732 | 0 | ref_000001 - ref_000732 |
reasoning |
1,056 | 1,056 | 0 | rea_000001 - rea_001056 |
OVEarth uses a two-level structure. Each entry in tasks.<task_key> references an item in annotations through ann_id.
{
"annotations": [
{
"id": "000001",
"image_path": "images/1.tif",
"width": 1225,
"height": 1225,
"category_name": "bell tower",
"segmentation": {
"size": [1225, 1225],
"counts": "aR\\R11VV12O1O1..."
},
"bbox": [[919, 676, 120, 39]],
"oriented_bbox": [
[[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]]
]
}
],
"tasks": {
"open_vocabulary": [
{"query_id": "ov_000001", "ann_id": "000001", "type": "positive", "phrase": "bell tower"},
{"query_id": "ov_000002", "ann_id": "000001", "type": "positive", "phrase": "church tower"},
{"query_id": "ov_000003", "ann_id": "000001", "type": "negative", "phrase": "water tower"}
],
"referring": [
{"query_id": "ref_000001", "ann_id": "000001", "query": "bell tower adjacent to the cemetery on the north side"}
],
"reasoning": [
{"query_id": "rea_000001", "ann_id": "000001", "query": "structure that houses large hanging bells used to call people for services"}
]
}
}For open_vocabulary, type is the only source of the positive/negative label:
positive: the queried target is present. The linked annotation supplies the ground-truth geometry.negative: the queried target is absent. The linked annotation may describe another object in the image and is not used as target geometry for that query.
All referring and reasoning queries are positive and store their text in query.
Ground-truth bbox values use COCO [x, y, width, height] format and are converted internally to [x1, y1, x2, y2]. Segmentation RLE is decoded directly.
The prediction file must be a UTF-8 JSON object. Its top-level tasks object groups predictions by task. Each evaluation run reads only the list selected by --task.
The counts strings below are shortened for readability. A real submission must contain complete, decodable COCO RLE data.
{
"tasks": {
"open_vocabulary": [
{
"query_id": "ov_000001",
"phrase": "bell tower",
"segmentation": {"size": [1225, 1225], "counts": "..."},
"bbox": [
{"bbox": [919, 676, 1039, 715], "score": 0.94}
],
"oriented_bbox": [
{
"polygon": [[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]],
"score": 0.91
}
]
},
{
"query_id": "ov_000003",
"phrase": "water tower",
"segmentation": null,
"bbox": [],
"oriented_bbox": []
}
],
"referring": [
{
"query_id": "ref_000001",
"query": "bell tower adjacent to the cemetery on the north side",
"segmentation": {"size": [1225, 1225], "counts": "..."},
"bbox": [],
"oriented_bbox": []
}
],
"reasoning": [
{
"query_id": "rea_000001",
"query": "structure that houses large hanging bells used to call people for services",
"segmentation": null,
"bbox": [[919, 676, 1039, 715]],
"oriented_bbox": []
}
]
}
}| JSON path | Type | Requirement | Description |
|---|---|---|---|
tasks |
object | Required for a standard submission | Task container; a missing value is treated as an empty object by the loader |
tasks.open_vocabulary |
list | Provide for open-vocabulary evaluation | Used by --task vocabulary; a missing list is treated as empty |
tasks.referring |
list | Provide for referring evaluation | Used by --task referring; a missing list is treated as empty |
tasks.reasoning |
list | Provide for reasoning evaluation | Used by --task reasoning; a missing list is treated as empty |
query_id |
string | Required in every prediction row | The only key used to match a prediction to a GT query |
phrase |
string | Optional | Readable metadata for open-vocabulary predictions; not used for matching |
query |
string | Optional | Readable metadata for referring/reasoning predictions; not used for matching |
segmentation |
object or null |
Recommended when segmentation is enabled | COCO RLE mask; null means no predicted mask |
bbox |
list | Recommended when bbox is enabled | Axis-aligned boxes; [] means no predicted boxes |
oriented_bbox |
list | Recommended when oriented bbox is enabled | Four-corner polygons; [] means no predicted boxes |
Fields outside the selected task or outside --modalities are ignored. A standard submission should include one row for every GT query_id. Omitting a known ID does not remove that sample from evaluation; it is scored as empty for every enabled modality. A prediction ID that is not present in the selected GT task causes an error.
segmentation uses COCO RLE. Both compressed strings and uncompressed integer lists are accepted:
{"size": [1225, 1225], "counts": "aR\\R11VV12O1O1..."}{"size": [2, 3], "counts": [2, 1, 3]}sizemust be[height, width]with non-negative integer values.- Except for the
[0, 0]empty-mask sentinel, the prediction size must exactly match the query's GT mask size, including for all-zero masks. countsmust be a valid compressed COCO string or a list of non-negative integers. Run lengths must sum toheight * width.- A single mask may contain at most 100,000,000 pixels.
null, a missing field, a valid all-zero RLE, and{"size": [0, 0], "counts": []}are treated as empty predictions.scoreand--score-thresholddo not apply to segmentation masks.
Each box may be an object with an optional confidence score or a bare coordinate array. The object form is recommended:
"bbox": [
{"bbox": [919, 676, 1039, 715], "score": 0.94},
{"bbox": [320, 180, 410, 265], "score": 0.81}
]Equivalent shorthand without scores:
"bbox": [
[919, 676, 1039, 715],
[320, 180, 410, 265]
]- Coordinates must be finite, absolute pixel values. Normalized coordinates in
[0, 1]are not supported. - With
--bbox-format xyxy(default), boxes are[x1, y1, x2, y2]and must satisfyx2 > x1andy2 > y1. - With
--bbox-format xywh, boxes are[x, y, width, height]and must have positive width and height. --bbox-formatapplies to the entire task file. Formats cannot be mixed per box.[]and a missing field mean no prediction.[0, 0, 0, 0]is accepted as a compatibility sentinel and ignored, but new submissions should use[].
Oriented boxes also support object and shorthand forms:
"oriented_bbox": [
{
"polygon": [[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]],
"score": 0.91
}
]"oriented_bbox": [
[[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]]
]polygonmust contain exactly four vertices, each represented by two finite absolute pixel coordinates[x, y].- Vertices must follow the polygon boundary in clockwise or counterclockwise order.
- The polygon must be convex and have positive area. Self-crossing vertex order is invalid.
[]and a missing field mean no prediction.--bbox-formatdoes not affect oriented boxes.
score applies only to bbox and oriented_bbox:
- Without
--score-threshold, scores may be omitted. Any provided score must be a finite numeric value. - With
--score-threshold T, every non-empty box must provide a score. Boxes withscore >= Tare retained; a missing score causes an error. - Confidence filtering occurs before IoU matching.
The required --modalities option declares the enabled modalities for one task evaluation. Valid values are segmentation, bbox, and oriented_bbox; separate multiple values with commas.
- Once a modality is enabled, every positive query with usable GT for that modality is evaluated.
null, an empty list, an all-zero mask, a missing field, or a missing prediction row is scored as an empty prediction. - A positive query without usable GT for an enabled modality is excluded from that modality's statistics.
- For a negative open-vocabulary query, any non-empty prediction in any enabled modality sets
pred_present = true. - Duplicate rows with the same
query_idare allowed. Equal-sized masks are merged by pixel union; bbox and oriented-bbox lists are concatenated. Different mask sizes cause an error. A standard submission should normally use one row per ID. --modalities autoinfers modalities from non-empty predictions and is intended only for debugging. It generally cannot infer a modality from an all-empty submission. Official benchmark runs must declare modalities explicitly.
ovearth-eval \
--gt ground_truth.json \
--predictions predictions.json \
--task vocabulary \
--modalities segmentation \
--output results.json \
--iou-threshold 0.5| Option | Type | Default | Description |
|---|---|---|---|
--gt |
path | Required | Ground-truth JSON path |
--prediction, --predictions |
path | Required | Prediction JSON path |
--task |
enum | vocabulary |
vocabulary, referring, or reasoning |
--output |
path | stdout | Output JSON path |
--iou-threshold |
float | 0.5 |
One-to-one bbox/oriented-bbox matching threshold; does not affect segmentation |
--score-threshold |
float | No filtering | Confidence threshold for bbox/oriented-bbox predictions only |
--bbox-format |
string | xyxy |
Prediction bbox format: xyxy or xywh |
--modalities |
comma-separated string | Required | Any combination of segmentation, bbox, and oriented_bbox; use auto only for exploration |
Examples:
# Open-vocabulary segmentation and bbox evaluation
ovearth-eval --gt gt.json --predictions preds.json --task vocabulary --modalities segmentation,bbox --output ov_results.json
# Referring-expression bbox evaluation
ovearth-eval --gt gt.json --predictions preds.json --task referring --modalities bbox --output ref_results.json
# Reasoning evaluation with oriented boxes
ovearth-eval --gt gt.json --predictions preds.json --task reasoning --modalities oriented_bbox --output rea_results.json
# Absolute-pixel xywh predictions with confidence filtering
ovearth-eval --gt gt.json --predictions preds.json --task vocabulary --modalities bbox --bbox-format xywh --score-threshold 0.3 --output ov_results.jsonRun the command once per task. A prediction file may contain all three task lists.
from ovearth_eval import evaluate_open_vocab, evaluate_referring, evaluate_reasoning
from ovearth_eval.io import (
load_open_vocab_samples,
load_referring_samples,
load_reasoning_samples,
)
samples = load_open_vocab_samples(
"gt.json",
"preds.json",
score_threshold=0.3,
bbox_format="xyxy",
modalities={"segmentation", "bbox"},
)
results = evaluate_open_vocab(samples, iou_threshold=0.5)
print(results["summary"]["presence"]["f1"])
print(results["summary"]["negative_accuracy"])
print(results["summary"]["seg"]["positive_micro_miou"])
print(results["summary"]["bbox"]["positive_micro_f1"])
print(results["summary"]["bbox"]["positive_mean_micro_f1@[0.5:0.95]"])
samples = load_referring_samples(
"gt.json",
"preds.json",
bbox_format="xyxy",
modalities={"bbox"},
)
results = evaluate_referring(samples, iou_threshold=0.5)
print(results["summary"]["bbox"]["positive_micro_f1"])
samples = load_reasoning_samples(
"gt.json",
"preds.json",
bbox_format="xyxy",
modalities={"oriented_bbox"},
)
results = evaluate_reasoning(samples, iou_threshold=0.5)
print(results["summary"]["oriented_bbox"]["positive_mean_micro_f1@[0.5:0.95]"])
for row in results["per_sample"]:
print(row["sample_id"], row["phrase"], row["oriented_bbox"])Each enabled summary dictionary contains a key_metrics object for convenient model comparison.
| Modality | Key metrics | Meaning |
|---|---|---|
presence |
mcc |
Matthews correlation coefficient for imbalanced presence labels |
seg |
positive_macro_precision, positive_macro_recall, positive_macro_miou, positive_micro_miou |
Pixel-level quality over positive queries |
bbox, oriented_bbox |
positive_micro_precision, positive_micro_recall, positive_micro_f1, positive_mean_micro_f1@[0.5:0.95] |
Detection quality over positive queries |
Presence metrics apply only to open_vocabulary. gt_present is determined by query type: positive is true and negative is false. After confidence filtering, any non-empty prediction in any enabled modality sets pred_present = true. An all-zero mask is empty.
| Metric | Formula | Meaning |
|---|---|---|
| TP | count | Positive query with a non-empty prediction |
| FP | count | Negative query with a non-empty prediction |
| TN | count | Negative query with an empty prediction |
| FN | count | Positive query with an empty prediction |
| Precision | TP / (TP + FP) | Fraction of predicted-present queries that are positive |
| Recall | TP / (TP + FN) | Fraction of positive queries predicted present |
| F1 | 2PR / (P + R) | Harmonic mean of precision and recall |
| Accuracy | (TP + TN) / N | Overall presence accuracy |
| FPR | FP / (FP + TN) | False-positive rate on negative queries |
| MCC | (TP*TN - FP*FN) / sqrt((TP+FP)(TP+FN)(TN+FP)(TN+FN)) |
Balanced correlation for imbalanced labels |
When a denominator is zero, the corresponding metric is null. Referring and reasoning contain only positive queries and do not report presence metrics.
Let I = |GT intersection Pred| and U = |GT union Pred|.
| Metric | Formula | Meaning |
|---|---|---|
| IoU | I / U | Per-query intersection over union |
| Dice / F1 | 2I / (` | GT |
| Pixel Precision | I / ` | Pred |
| Pixel Recall | I / ` | GT |
positive_micro_miou |
sum(I) / sum(U) | Micro IoU over positive queries |
positive_macro_miou |
mean(per-query IoU) | Equal-weight mean IoU over positive queries |
Only positive queries contribute to segmentation quality summaries. Negative open-vocabulary behavior is measured separately by negative_false_positive_rate, preventing empty-empty pairs from inflating IoU or Dice.
Axis-aligned and oriented detection use thresholded maximum-cardinality one-to-one IoU matching. Each GT and prediction can participate in at most one match. Matching first maximizes the number of valid pairs, avoiding the under-counting that can occur with a locally greedy highest-IoU strategy.
| Metric | Meaning |
|---|---|
| TP | Number of one-to-one pairs with IoU at or above the threshold |
| FP | Number of unmatched predictions |
| FN | Number of unmatched GT boxes |
positive_micro_precision |
Global TP / (TP + FP) |
positive_micro_recall |
Global TP / (TP + FN) |
positive_micro_f1 |
F1 computed from global TP, FP, and FN |
positive_macro_f1 |
Mean per-positive-query F1 |
positive_macro_matched_iou |
Mean IoU across matched pairs |
positive_mean_micro_f1@[0.5:0.95] |
Mean micro F1 at IoU thresholds 0.50, 0.55, ..., 0.95 |
{
"task": "open_vocabulary",
"config": {
"iou_threshold": 0.5,
"modalities": ["bbox", "segmentation"]
},
"summary": {
"num_samples": 100,
"num_positive": 80,
"num_negative": 20,
"negative_accuracy": 0.85,
"negative_false_positive_rate": 0.15,
"presence": {
"tp": 72,
"fp": 3,
"tn": 17,
"fn": 8,
"precision": 0.96,
"recall": 0.90,
"f1": 0.93,
"accuracy": 0.89,
"fpr": 0.15,
"mcc": 0.82,
"key_metrics": {"mcc": 0.82}
},
"seg": {
"num_samples_evaluated": 80,
"num_positive_evaluated": 80,
"positive_micro_miou": 0.7142,
"positive_macro_miou": 0.7201,
"positive_macro_dice": 0.8031,
"positive_macro_precision": 0.8512,
"positive_macro_recall": 0.7698,
"key_metrics": {
"positive_macro_precision": 0.8512,
"positive_macro_recall": 0.7698,
"positive_macro_miou": 0.7201,
"positive_micro_miou": 0.7142
}
},
"bbox": {
"num_samples_evaluated": 65,
"num_positive_evaluated": 65,
"tp": 120,
"fp": 15,
"fn": 30,
"positive_micro_precision": 0.8889,
"positive_micro_recall": 0.8000,
"positive_micro_f1": 0.8421,
"positive_macro_f1": 0.8134,
"positive_macro_matched_iou": 0.7231,
"positive_mean_micro_f1@[0.5:0.95]": 0.7512,
"key_metrics": {
"positive_micro_precision": 0.8889,
"positive_micro_recall": 0.8000,
"positive_micro_f1": 0.8421,
"positive_mean_micro_f1@[0.5:0.95]": 0.7512
}
},
"oriented_bbox": null
},
"per_sample": [
{
"sample_id": "ov_000001",
"phrase": "bell tower",
"negative": false,
"gt_present": true,
"pred_present": true,
"image_path": "images/1.tif",
"seg": {
"gt_area": 4680,
"pred_area": 4512,
"intersection": 4201,
"union": 4991,
"iou": 0.7153,
"dice": 0.8342,
"precision": 0.8522,
"recall": 0.8167,
"f1": 0.8342,
"gt_present": true,
"pred_present": true
},
"bbox": {
"num_gt": 1,
"num_pred": 1,
"tp": 1,
"fp": 0,
"fn": 0,
"precision": 1.0,
"recall": 1.0,
"f1": 1.0,
"mean_matched_iou": 0.7812,
"gt_present": true,
"pred_present": true
},
"oriented_bbox": null
}
]
}When a modality is disabled, its summary and per-sample values are null.
These tasks use the same modality summaries and per-sample schema. Their summary objects do not contain num_negative, negative_accuracy, negative_false_positive_rate, or presence. Every per-sample negative value is false.
The following abbreviated example omits most summary values and per-sample rows:
{
"task": "referring",
"config": {
"iou_threshold": 0.5,
"modalities": ["segmentation"]
},
"summary": {
"num_samples": 50,
"seg": {
"num_samples_evaluated": 50,
"num_positive_evaluated": 50,
"positive_micro_miou": 0.71,
"positive_macro_miou": 0.68,
"positive_macro_dice": 0.76,
"positive_macro_precision": 0.79,
"positive_macro_recall": 0.74,
"key_metrics": {
"positive_macro_precision": 0.79,
"positive_macro_recall": 0.74,
"positive_macro_miou": 0.68,
"positive_micro_miou": 0.71
}
},
"bbox": null,
"oriented_bbox": null
},
"per_sample": [
{
"sample_id": "ref_000001",
"phrase": "bell tower adjacent to the cemetery on the north side",
"negative": false,
"gt_present": true,
"pred_present": true,
"image_path": "images/1.tif",
"seg": {
"gt_area": 4680,
"pred_area": 4512,
"intersection": 4201,
"union": 4991,
"iou": 0.7153,
"dice": 0.8342,
"precision": 0.8522,
"recall": 0.8167,
"f1": 0.8342,
"gt_present": true,
"pred_present": true
},
"bbox": null,
"oriented_bbox": null
}
]
}This section applies only to open_vocabulary.
A row with tasks.open_vocabulary[i].type == "negative" is a negative query. type is the only label source. The linked annotation may contain another category and is not used for localization quality on the negative query.
The correct output for a negative query is empty in every enabled modality:
- Segmentation:
segmentation == nullor an all-zero mask - Axis-aligned boxes:
bbox == [] - Oriented boxes:
oriented_bbox == []
Negative empty-empty pairs do not contribute to positive-query IoU or Dice averages. Hallucinations are measured through presence metrics and negative_false_positive_rate.
An oriented bounding box is represented by a four-corner polygon:
"oriented_bbox": [
{"polygon": [[14, 21], [79, 23], [77, 69], [12, 67]]}
]Polygon IoU is computed with Sutherland-Hodgman clipping:
- Use each edge of one polygon as a clipping boundary.
- Repeatedly clip the other polygon to the inside of that edge.
- Compute the intersection area with the shoelace formula.
- Compute
IoU = intersection / (area_a + area_b - intersection).
Vertices may be clockwise or counterclockwise, but they must follow the polygon boundary. Self-crossing order, non-convex quadrilaterals, and zero-area polygons are rejected.
| Property | bbox |
oriented_bbox |
|---|---|---|
| Representation | [x1, y1, x2, y2] by default |
Four boundary-ordered vertices |
| IoU | Axis-aligned rectangle intersection | Convex polygon clipping |
| Typical use | Axis-aligned targets | Targets at arbitrary orientations |
| GT conversion | COCO xywh converted to xyxy | No coordinate-format conversion |
pip install -e ".[dev]"
pytest -q
pytest tests/test_segmentation.py -v
pytest tests/test_detection.py -v
pytest tests/test_polygon.py -v
pytest tests/test_open_vocab.py -v
pytest tests/test_io.py -v
pytest tests/test_cli.py -v| Test file | Coverage |
|---|---|
test_presence.py |
TP/FP/TN/FN, MCC edge cases, precision, recall, and F1 |
test_segmentation.py |
Exact overlap and empty-mask behavior |
test_detection.py |
One-to-one matching, duplicate predictions, coordinate conversion, invalid boxes |
test_polygon.py |
Oriented IoU and polygon-clipping edge cases |
test_rle.py |
COCO RLE validation and round trips |
test_open_vocab.py |
Multi-modality evaluation, negative queries, referring and reasoning paths |
test_io.py |
JSON loading and confidence filtering |
test_cli.py |
End-to-end CLI smoke tests |
| Package | Version | Purpose |
|---|---|---|
numpy |
>=1.23 |
Mask operations and metric aggregation |
Pillow |
>=9.0 |
Image-format support |
typer |
>=0.9 |
Command-line interface |
| Package | Purpose |
|---|---|
pytest |
Test framework |
All dependencies are declared in pyproject.toml and installed by pip install ..