Reading Order Independent Metrics for Information Extraction in Handwritten Documents


手書き文書の情報抽出プロセスは、自動転写の取得と、その転写に対する固有表現認識 (NER) の実行に依存する傾向があります。
私たちの実験では、メトリクスの動作を詳細に分析して、タスクを正しく評価するための最小のメトリクス セットと考えられるものを推奨します。


Information Extraction processes in handwritten documents tend to rely on obtaining an automatic transcription and performing Named Entity Recognition (NER) over such transcription. For this reason, in publicly available datasets, the performance of the systems is usually evaluated with metrics particular to each dataset. Moreover, most of the metrics employed are sensitive to reading order errors. Therefore, they do not reflect the expected final application of the system and introduce biases in more complex documents. In this paper, we propose and publicly release a set of reading order independent metrics tailored to Information Extraction evaluation in handwritten documents. In our experimentation, we perform an in-depth analysis of the behavior of the metrics to recommend what we consider to be the minimal set of metrics to evaluate a task correctly.


著者 David Villanova-Aparisi,Solène Tarride,Carlos-D. Martínez-Hinarejos,Verónica Romero,Christopher Kermorvant,Moisés Pastor-Gadea
発行日 2024-04-29 12:49:30+00:00
arxivサイト arxiv_id(pdf)

提供元, 利用サービス, Google

カテゴリー: cs.CV パーマリンク