Identity Overlap Between Face Recognition Train/Test Data: Causing Optimistic Bias in Accuracy Measurement

要約

パターン認識の基本原則は、トレーニングセットとテストセット間の重複により楽観的な精度推定が生じるということです。
顔認識用のディープ CNN は、トレーニングセット内のアイデンティティの N 方向分類用にトレーニングされます。
精度は通常、LFW、CALFW、CPLFW、CFP-FP、AgeDB-30 などのテストセットからの画像ペアの平均 10 倍の分類精度として推定されます。
トレーニングセットとテストセットは独立して組み立てられているため、任意のテストセット内の画像と ID が任意のトレーニングセットにも存在する可能性があります。
特に、私たちの実験では、LFW ファミリのテストセットと MS1MV2 トレーニングセットの間で驚くべき程度の同一性とイメージの重複が明らかになりました。
私たちの実験では、MS1MV2 の識別ラベルノイズも明らかになりました。
楽観的バイアスの大きさを明らかにするために、同一サイズで同一サイズの MS1MV2 サブセットと同一サイズで同一でないサブセットで達成される精度を LFW と比較します。
LFW ファミリーのより困難なテストセットを使用すると、より困難なテストセットほど楽観的バイアスのサイズが大きくなることがわかります。
私たちの結果は、顔認識研究におけるアイデンティティに素なトレーニングとテスト方法論の欠如とその必要性を浮き彫りにしています。

要約(オリジナル)

A fundamental tenet of pattern recognition is that overlap between training and testing sets causes an optimistic accuracy estimate. Deep CNNs for face recognition are trained for N-way classification of the identities in the training set. Accuracy is commonly estimated as average 10-fold classification accuracy on image pairs from test sets such as LFW, CALFW, CPLFW, CFP-FP and AgeDB-30. Because train and test sets have been independently assembled, images and identities in any given test set may also be present in any given training set. In particular, our experiments reveal a surprising degree of identity and image overlap between the LFW family of test sets and the MS1MV2 training set. Our experiments also reveal identity label noise in MS1MV2. We compare accuracy achieved with same-size MS1MV2 subsets that are identity-disjoint and not identity-disjoint with LFW, to reveal the size of the optimistic bias. Using more challenging test sets from the LFW family, we find that the size of the optimistic bias is larger for more challenging test sets. Our results highlight the lack of and the need for identity-disjoint train and test methodology in face recognition research.

arxiv情報

著者	Haiyu Wu,Sicong Tian,Jacob Gutierrez,Aman Bhatta,Kağan Öztürk,Kevin W. Bowyer
発行日	2024-05-15 14:59:26+00:00
arxivサイト	arxiv_id(pdf)

提供元, 利用サービス

arxiv.jp, Google

Identity Overlap Between Face Recognition Train/Test Data: Causing Optimistic Bias in Accuracy Measurement

要約

要約(オリジナル)

arxiv情報

提供元, 利用サービス

最近の投稿

最近のコメント

アーカイブ

カテゴリー