Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

OVEarth-Bench Evaluation Toolkit

English | 中文

OVEarth-Bench is a Python evaluation suite for remote-sensing image understanding. It supports three tasks: open-vocabulary segmentation/detection (open_vocabulary), referring-expression segmentation/grounding (referring), and reasoning segmentation/grounding (reasoning).


Table of Contents

  1. Installation
  2. Project Layout
  3. Tasks
  4. Data Formats
  5. Command-Line Interface
  6. Python API
  7. Metrics
  8. Output JSON
  9. Negative Queries
  10. Oriented Bounding Box Evaluation
  11. Tests
  12. Dependencies

Installation

git clone <repo-url>
cd OVEarth-bench

# Standard installation
pip install .

# Editable installation with test dependencies
pip install -e ".[dev]"

After installation, the ovearth-eval command is available:

ovearth-eval --help

Project Layout

(project root)/
|-- pyproject.toml
|-- README.md
|-- README-ZH.md
|-- ovearth_eval/
|   |-- __init__.py
|   |-- _utils.py
|   |-- types.py
|   |-- presence.py
|   |-- segmentation.py
|   |-- detection.py
|   |-- oriented_detection.py
|   |-- polygon.py
|   |-- rle.py
|   |-- open_vocab.py
|   |-- io.py
|   `-- cli.py
`-- tests/
    |-- test_presence.py
    |-- test_segmentation.py
    |-- test_detection.py
    |-- test_polygon.py
    |-- test_rle.py
    |-- test_open_vocab.py
    |-- test_io.py
    `-- test_cli.py

Tasks

Select a task with --task:

Task CLI value Prediction list Samples Task-specific metrics
Open-vocabulary segmentation/detection vocabulary tasks.open_vocabulary Positive and negative Presence metrics and negative-query false-positive rate
Referring-expression segmentation/localization referring tasks.referring Positive only None
Reasoning segmentation/localization reasoning tasks.reasoning Positive only None
  • Open vocabulary: Given a category phrase such as "bell tower", the model must determine whether that category is present and then segment or localize it. Negative queries measure hallucination suppression.
  • Referring: Given a spatial or contextual description, the model must segment or localize the referred object.
  • Reasoning: Given a functional or attribute-based description, the model must infer the target and segment or localize it.

The current ground_truth.json contains 590 annotations, 172 categories, and 5,023 queries:

Task Queries Positive Negative query_id range
open_vocabulary 3,235 1,067 2,168 ov_000001 - ov_003235
referring 732 732 0 ref_000001 - ref_000732
reasoning 1,056 1,056 0 rea_000001 - rea_001056

Data Formats

Ground Truth

OVEarth uses a two-level structure. Each entry in tasks.<task_key> references an item in annotations through ann_id.

{
  "annotations": [
    {
      "id": "000001",
      "image_path": "images/1.tif",
      "width": 1225,
      "height": 1225,
      "category_name": "bell tower",
      "segmentation": {
        "size": [1225, 1225],
        "counts": "aR\\R11VV12O1O1..."
      },
      "bbox": [[919, 676, 120, 39]],
      "oriented_bbox": [
        [[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]]
      ]
    }
  ],
  "tasks": {
    "open_vocabulary": [
      {"query_id": "ov_000001", "ann_id": "000001", "type": "positive", "phrase": "bell tower"},
      {"query_id": "ov_000002", "ann_id": "000001", "type": "positive", "phrase": "church tower"},
      {"query_id": "ov_000003", "ann_id": "000001", "type": "negative", "phrase": "water tower"}
    ],
    "referring": [
      {"query_id": "ref_000001", "ann_id": "000001", "query": "bell tower adjacent to the cemetery on the north side"}
    ],
    "reasoning": [
      {"query_id": "rea_000001", "ann_id": "000001", "query": "structure that houses large hanging bells used to call people for services"}
    ]
  }
}

For open_vocabulary, type is the only source of the positive/negative label:

  • positive: the queried target is present. The linked annotation supplies the ground-truth geometry.
  • negative: the queried target is absent. The linked annotation may describe another object in the image and is not used as target geometry for that query.

All referring and reasoning queries are positive and store their text in query.

Ground-truth bbox values use COCO [x, y, width, height] format and are converted internally to [x1, y1, x2, y2]. Segmentation RLE is decoded directly.

Prediction File

The prediction file must be a UTF-8 JSON object. Its top-level tasks object groups predictions by task. Each evaluation run reads only the list selected by --task.

Complete Example

The counts strings below are shortened for readability. A real submission must contain complete, decodable COCO RLE data.

{
  "tasks": {
    "open_vocabulary": [
      {
        "query_id": "ov_000001",
        "phrase": "bell tower",
        "segmentation": {"size": [1225, 1225], "counts": "..."},
        "bbox": [
          {"bbox": [919, 676, 1039, 715], "score": 0.94}
        ],
        "oriented_bbox": [
          {
            "polygon": [[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]],
            "score": 0.91
          }
        ]
      },
      {
        "query_id": "ov_000003",
        "phrase": "water tower",
        "segmentation": null,
        "bbox": [],
        "oriented_bbox": []
      }
    ],
    "referring": [
      {
        "query_id": "ref_000001",
        "query": "bell tower adjacent to the cemetery on the north side",
        "segmentation": {"size": [1225, 1225], "counts": "..."},
        "bbox": [],
        "oriented_bbox": []
      }
    ],
    "reasoning": [
      {
        "query_id": "rea_000001",
        "query": "structure that houses large hanging bells used to call people for services",
        "segmentation": null,
        "bbox": [[919, 676, 1039, 715]],
        "oriented_bbox": []
      }
    ]
  }
}

Top-Level and Common Fields

JSON path Type Requirement Description
tasks object Required for a standard submission Task container; a missing value is treated as an empty object by the loader
tasks.open_vocabulary list Provide for open-vocabulary evaluation Used by --task vocabulary; a missing list is treated as empty
tasks.referring list Provide for referring evaluation Used by --task referring; a missing list is treated as empty
tasks.reasoning list Provide for reasoning evaluation Used by --task reasoning; a missing list is treated as empty
query_id string Required in every prediction row The only key used to match a prediction to a GT query
phrase string Optional Readable metadata for open-vocabulary predictions; not used for matching
query string Optional Readable metadata for referring/reasoning predictions; not used for matching
segmentation object or null Recommended when segmentation is enabled COCO RLE mask; null means no predicted mask
bbox list Recommended when bbox is enabled Axis-aligned boxes; [] means no predicted boxes
oriented_bbox list Recommended when oriented bbox is enabled Four-corner polygons; [] means no predicted boxes

Fields outside the selected task or outside --modalities are ignored. A standard submission should include one row for every GT query_id. Omitting a known ID does not remove that sample from evaluation; it is scored as empty for every enabled modality. A prediction ID that is not present in the selected GT task causes an error.

Segmentation

segmentation uses COCO RLE. Both compressed strings and uncompressed integer lists are accepted:

{"size": [1225, 1225], "counts": "aR\\R11VV12O1O1..."}
{"size": [2, 3], "counts": [2, 1, 3]}
  • size must be [height, width] with non-negative integer values.
  • Except for the [0, 0] empty-mask sentinel, the prediction size must exactly match the query's GT mask size, including for all-zero masks.
  • counts must be a valid compressed COCO string or a list of non-negative integers. Run lengths must sum to height * width.
  • A single mask may contain at most 100,000,000 pixels.
  • null, a missing field, a valid all-zero RLE, and {"size": [0, 0], "counts": []} are treated as empty predictions.
  • score and --score-threshold do not apply to segmentation masks.

Axis-Aligned Bounding Boxes

Each box may be an object with an optional confidence score or a bare coordinate array. The object form is recommended:

"bbox": [
  {"bbox": [919, 676, 1039, 715], "score": 0.94},
  {"bbox": [320, 180, 410, 265], "score": 0.81}
]

Equivalent shorthand without scores:

"bbox": [
  [919, 676, 1039, 715],
  [320, 180, 410, 265]
]
  • Coordinates must be finite, absolute pixel values. Normalized coordinates in [0, 1] are not supported.
  • With --bbox-format xyxy (default), boxes are [x1, y1, x2, y2] and must satisfy x2 > x1 and y2 > y1.
  • With --bbox-format xywh, boxes are [x, y, width, height] and must have positive width and height.
  • --bbox-format applies to the entire task file. Formats cannot be mixed per box.
  • [] and a missing field mean no prediction. [0, 0, 0, 0] is accepted as a compatibility sentinel and ignored, but new submissions should use [].

Oriented Bounding Boxes

Oriented boxes also support object and shorthand forms:

"oriented_bbox": [
  {
    "polygon": [[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]],
    "score": 0.91
  }
]
"oriented_bbox": [
  [[1036.4, 714.39], [919.0, 706.0], [921.29, 673.88], [1038.7, 682.26]]
]
  • polygon must contain exactly four vertices, each represented by two finite absolute pixel coordinates [x, y].
  • Vertices must follow the polygon boundary in clockwise or counterclockwise order.
  • The polygon must be convex and have positive area. Self-crossing vertex order is invalid.
  • [] and a missing field mean no prediction.
  • --bbox-format does not affect oriented boxes.

Confidence Filtering

score applies only to bbox and oriented_bbox:

  • Without --score-threshold, scores may be omitted. Any provided score must be a finite numeric value.
  • With --score-threshold T, every non-empty box must provide a score. Boxes with score >= T are retained; a missing score causes an error.
  • Confidence filtering occurs before IoU matching.

Modalities, Empty Predictions, and Duplicate IDs

The required --modalities option declares the enabled modalities for one task evaluation. Valid values are segmentation, bbox, and oriented_bbox; separate multiple values with commas.

  • Once a modality is enabled, every positive query with usable GT for that modality is evaluated. null, an empty list, an all-zero mask, a missing field, or a missing prediction row is scored as an empty prediction.
  • A positive query without usable GT for an enabled modality is excluded from that modality's statistics.
  • For a negative open-vocabulary query, any non-empty prediction in any enabled modality sets pred_present = true.
  • Duplicate rows with the same query_id are allowed. Equal-sized masks are merged by pixel union; bbox and oriented-bbox lists are concatenated. Different mask sizes cause an error. A standard submission should normally use one row per ID.
  • --modalities auto infers modalities from non-empty predictions and is intended only for debugging. It generally cannot infer a modality from an all-empty submission. Official benchmark runs must declare modalities explicitly.

Command-Line Interface

Basic Usage

ovearth-eval \
  --gt ground_truth.json \
  --predictions predictions.json \
  --task vocabulary \
  --modalities segmentation \
  --output results.json \
  --iou-threshold 0.5
Option Type Default Description
--gt path Required Ground-truth JSON path
--prediction, --predictions path Required Prediction JSON path
--task enum vocabulary vocabulary, referring, or reasoning
--output path stdout Output JSON path
--iou-threshold float 0.5 One-to-one bbox/oriented-bbox matching threshold; does not affect segmentation
--score-threshold float No filtering Confidence threshold for bbox/oriented-bbox predictions only
--bbox-format string xyxy Prediction bbox format: xyxy or xywh
--modalities comma-separated string Required Any combination of segmentation, bbox, and oriented_bbox; use auto only for exploration

Examples:

# Open-vocabulary segmentation and bbox evaluation
ovearth-eval --gt gt.json --predictions preds.json --task vocabulary --modalities segmentation,bbox --output ov_results.json

# Referring-expression bbox evaluation
ovearth-eval --gt gt.json --predictions preds.json --task referring --modalities bbox --output ref_results.json

# Reasoning evaluation with oriented boxes
ovearth-eval --gt gt.json --predictions preds.json --task reasoning --modalities oriented_bbox --output rea_results.json

# Absolute-pixel xywh predictions with confidence filtering
ovearth-eval --gt gt.json --predictions preds.json --task vocabulary --modalities bbox --bbox-format xywh --score-threshold 0.3 --output ov_results.json

Run the command once per task. A prediction file may contain all three task lists.


Python API

from ovearth_eval import evaluate_open_vocab, evaluate_referring, evaluate_reasoning
from ovearth_eval.io import (
    load_open_vocab_samples,
    load_referring_samples,
    load_reasoning_samples,
)

samples = load_open_vocab_samples(
    "gt.json",
    "preds.json",
    score_threshold=0.3,
    bbox_format="xyxy",
    modalities={"segmentation", "bbox"},
)
results = evaluate_open_vocab(samples, iou_threshold=0.5)
print(results["summary"]["presence"]["f1"])
print(results["summary"]["negative_accuracy"])
print(results["summary"]["seg"]["positive_micro_miou"])
print(results["summary"]["bbox"]["positive_micro_f1"])
print(results["summary"]["bbox"]["positive_mean_micro_f1@[0.5:0.95]"])

samples = load_referring_samples(
    "gt.json",
    "preds.json",
    bbox_format="xyxy",
    modalities={"bbox"},
)
results = evaluate_referring(samples, iou_threshold=0.5)
print(results["summary"]["bbox"]["positive_micro_f1"])

samples = load_reasoning_samples(
    "gt.json",
    "preds.json",
    bbox_format="xyxy",
    modalities={"oriented_bbox"},
)
results = evaluate_reasoning(samples, iou_threshold=0.5)
print(results["summary"]["oriented_bbox"]["positive_mean_micro_f1@[0.5:0.95]"])

for row in results["per_sample"]:
    print(row["sample_id"], row["phrase"], row["oriented_bbox"])

Metrics

Key Metrics

Each enabled summary dictionary contains a key_metrics object for convenient model comparison.

Modality Key metrics Meaning
presence mcc Matthews correlation coefficient for imbalanced presence labels
seg positive_macro_precision, positive_macro_recall, positive_macro_miou, positive_micro_miou Pixel-level quality over positive queries
bbox, oriented_bbox positive_micro_precision, positive_micro_recall, positive_micro_f1, positive_mean_micro_f1@[0.5:0.95] Detection quality over positive queries

Presence Metrics

Presence metrics apply only to open_vocabulary. gt_present is determined by query type: positive is true and negative is false. After confidence filtering, any non-empty prediction in any enabled modality sets pred_present = true. An all-zero mask is empty.

Metric Formula Meaning
TP count Positive query with a non-empty prediction
FP count Negative query with a non-empty prediction
TN count Negative query with an empty prediction
FN count Positive query with an empty prediction
Precision TP / (TP + FP) Fraction of predicted-present queries that are positive
Recall TP / (TP + FN) Fraction of positive queries predicted present
F1 2PR / (P + R) Harmonic mean of precision and recall
Accuracy (TP + TN) / N Overall presence accuracy
FPR FP / (FP + TN) False-positive rate on negative queries
MCC (TP*TN - FP*FN) / sqrt((TP+FP)(TP+FN)(TN+FP)(TN+FN)) Balanced correlation for imbalanced labels

When a denominator is zero, the corresponding metric is null. Referring and reasoning contain only positive queries and do not report presence metrics.

Segmentation Metrics

Let I = |GT intersection Pred| and U = |GT union Pred|.

Metric Formula Meaning
IoU I / U Per-query intersection over union
Dice / F1 2I / (` GT
Pixel Precision I / ` Pred
Pixel Recall I / ` GT
positive_micro_miou sum(I) / sum(U) Micro IoU over positive queries
positive_macro_miou mean(per-query IoU) Equal-weight mean IoU over positive queries

Only positive queries contribute to segmentation quality summaries. Negative open-vocabulary behavior is measured separately by negative_false_positive_rate, preventing empty-empty pairs from inflating IoU or Dice.

Detection Metrics

Axis-aligned and oriented detection use thresholded maximum-cardinality one-to-one IoU matching. Each GT and prediction can participate in at most one match. Matching first maximizes the number of valid pairs, avoiding the under-counting that can occur with a locally greedy highest-IoU strategy.

Metric Meaning
TP Number of one-to-one pairs with IoU at or above the threshold
FP Number of unmatched predictions
FN Number of unmatched GT boxes
positive_micro_precision Global TP / (TP + FP)
positive_micro_recall Global TP / (TP + FN)
positive_micro_f1 F1 computed from global TP, FP, and FN
positive_macro_f1 Mean per-positive-query F1
positive_macro_matched_iou Mean IoU across matched pairs
positive_mean_micro_f1@[0.5:0.95] Mean micro F1 at IoU thresholds 0.50, 0.55, ..., 0.95

Output JSON

Open-Vocabulary Output

{
  "task": "open_vocabulary",
  "config": {
    "iou_threshold": 0.5,
    "modalities": ["bbox", "segmentation"]
  },
  "summary": {
    "num_samples": 100,
    "num_positive": 80,
    "num_negative": 20,
    "negative_accuracy": 0.85,
    "negative_false_positive_rate": 0.15,
    "presence": {
      "tp": 72,
      "fp": 3,
      "tn": 17,
      "fn": 8,
      "precision": 0.96,
      "recall": 0.90,
      "f1": 0.93,
      "accuracy": 0.89,
      "fpr": 0.15,
      "mcc": 0.82,
      "key_metrics": {"mcc": 0.82}
    },
    "seg": {
      "num_samples_evaluated": 80,
      "num_positive_evaluated": 80,
      "positive_micro_miou": 0.7142,
      "positive_macro_miou": 0.7201,
      "positive_macro_dice": 0.8031,
      "positive_macro_precision": 0.8512,
      "positive_macro_recall": 0.7698,
      "key_metrics": {
        "positive_macro_precision": 0.8512,
        "positive_macro_recall": 0.7698,
        "positive_macro_miou": 0.7201,
        "positive_micro_miou": 0.7142
      }
    },
    "bbox": {
      "num_samples_evaluated": 65,
      "num_positive_evaluated": 65,
      "tp": 120,
      "fp": 15,
      "fn": 30,
      "positive_micro_precision": 0.8889,
      "positive_micro_recall": 0.8000,
      "positive_micro_f1": 0.8421,
      "positive_macro_f1": 0.8134,
      "positive_macro_matched_iou": 0.7231,
      "positive_mean_micro_f1@[0.5:0.95]": 0.7512,
      "key_metrics": {
        "positive_micro_precision": 0.8889,
        "positive_micro_recall": 0.8000,
        "positive_micro_f1": 0.8421,
        "positive_mean_micro_f1@[0.5:0.95]": 0.7512
      }
    },
    "oriented_bbox": null
  },
  "per_sample": [
    {
      "sample_id": "ov_000001",
      "phrase": "bell tower",
      "negative": false,
      "gt_present": true,
      "pred_present": true,
      "image_path": "images/1.tif",
      "seg": {
        "gt_area": 4680,
        "pred_area": 4512,
        "intersection": 4201,
        "union": 4991,
        "iou": 0.7153,
        "dice": 0.8342,
        "precision": 0.8522,
        "recall": 0.8167,
        "f1": 0.8342,
        "gt_present": true,
        "pred_present": true
      },
      "bbox": {
        "num_gt": 1,
        "num_pred": 1,
        "tp": 1,
        "fp": 0,
        "fn": 0,
        "precision": 1.0,
        "recall": 1.0,
        "f1": 1.0,
        "mean_matched_iou": 0.7812,
        "gt_present": true,
        "pred_present": true
      },
      "oriented_bbox": null
    }
  ]
}

When a modality is disabled, its summary and per-sample values are null.

Referring and Reasoning Output

These tasks use the same modality summaries and per-sample schema. Their summary objects do not contain num_negative, negative_accuracy, negative_false_positive_rate, or presence. Every per-sample negative value is false.

The following abbreviated example omits most summary values and per-sample rows:

{
  "task": "referring",
  "config": {
    "iou_threshold": 0.5,
    "modalities": ["segmentation"]
  },
  "summary": {
    "num_samples": 50,
    "seg": {
      "num_samples_evaluated": 50,
      "num_positive_evaluated": 50,
      "positive_micro_miou": 0.71,
      "positive_macro_miou": 0.68,
      "positive_macro_dice": 0.76,
      "positive_macro_precision": 0.79,
      "positive_macro_recall": 0.74,
      "key_metrics": {
        "positive_macro_precision": 0.79,
        "positive_macro_recall": 0.74,
        "positive_macro_miou": 0.68,
        "positive_micro_miou": 0.71
      }
    },
    "bbox": null,
    "oriented_bbox": null
  },
  "per_sample": [
    {
      "sample_id": "ref_000001",
      "phrase": "bell tower adjacent to the cemetery on the north side",
      "negative": false,
      "gt_present": true,
      "pred_present": true,
      "image_path": "images/1.tif",
      "seg": {
        "gt_area": 4680,
        "pred_area": 4512,
        "intersection": 4201,
        "union": 4991,
        "iou": 0.7153,
        "dice": 0.8342,
        "precision": 0.8522,
        "recall": 0.8167,
        "f1": 0.8342,
        "gt_present": true,
        "pred_present": true
      },
      "bbox": null,
      "oriented_bbox": null
    }
  ]
}

Negative Queries

This section applies only to open_vocabulary.

A row with tasks.open_vocabulary[i].type == "negative" is a negative query. type is the only label source. The linked annotation may contain another category and is not used for localization quality on the negative query.

The correct output for a negative query is empty in every enabled modality:

  • Segmentation: segmentation == null or an all-zero mask
  • Axis-aligned boxes: bbox == []
  • Oriented boxes: oriented_bbox == []

Negative empty-empty pairs do not contribute to positive-query IoU or Dice averages. Hallucinations are measured through presence metrics and negative_false_positive_rate.


Oriented Bounding Box Evaluation

An oriented bounding box is represented by a four-corner polygon:

"oriented_bbox": [
  {"polygon": [[14, 21], [79, 23], [77, 69], [12, 67]]}
]

Polygon IoU is computed with Sutherland-Hodgman clipping:

  1. Use each edge of one polygon as a clipping boundary.
  2. Repeatedly clip the other polygon to the inside of that edge.
  3. Compute the intersection area with the shoelace formula.
  4. Compute IoU = intersection / (area_a + area_b - intersection).

Vertices may be clockwise or counterclockwise, but they must follow the polygon boundary. Self-crossing order, non-convex quadrilaterals, and zero-area polygons are rejected.

Property bbox oriented_bbox
Representation [x1, y1, x2, y2] by default Four boundary-ordered vertices
IoU Axis-aligned rectangle intersection Convex polygon clipping
Typical use Axis-aligned targets Targets at arbitrary orientations
GT conversion COCO xywh converted to xyxy No coordinate-format conversion

Tests

pip install -e ".[dev]"
pytest -q

pytest tests/test_segmentation.py -v
pytest tests/test_detection.py -v
pytest tests/test_polygon.py -v
pytest tests/test_open_vocab.py -v
pytest tests/test_io.py -v
pytest tests/test_cli.py -v
Test file Coverage
test_presence.py TP/FP/TN/FN, MCC edge cases, precision, recall, and F1
test_segmentation.py Exact overlap and empty-mask behavior
test_detection.py One-to-one matching, duplicate predictions, coordinate conversion, invalid boxes
test_polygon.py Oriented IoU and polygon-clipping edge cases
test_rle.py COCO RLE validation and round trips
test_open_vocab.py Multi-modality evaluation, negative queries, referring and reasoning paths
test_io.py JSON loading and confidence filtering
test_cli.py End-to-end CLI smoke tests

Dependencies

Runtime

Package Version Purpose
numpy >=1.23 Mask operations and metric aggregation
Pillow >=9.0 Image-format support
typer >=0.9 Command-line interface

Development

Package Purpose
pytest Test framework

All dependencies are declared in pyproject.toml and installed by pip install ..

Releases

Packages

Contributors

Languages