Open Access

Table D.1

Perfect match rates by LLM, query generation method, query difficulty, and with or without self-correction.

Simple Medium Hard
Model w/o Self-Corr w/ Self-Corr w/o Self-Corr w/ Self-Corr w/o Self-Corr w/ Self-Corr
ID Match (Rows) - Direct
Claude 3.7 0.844 ± 0.0 0.866 ± 0.021 0.160 ± 0.084 0.160 ± 0.084 0.280 ± 0.079 0.380 ± 0.079
Claude Opus 4.6 0.844 ± 0.0 0.938 ± 0.0 0.250 ± 0.053 0.400 ± 0.047 0.400 ± 0.0 0.480 ± 0.042
Claude Sonnet 4.5 0.844 ± 0.0 0.875 ± 0.0 0.400 ± 0.0 0.500 ± 0.0 0.200 ± 0.0 0.200 ± 0.0
Gemini 2.5 Flash 0.775 ± 0.020 0.819 ± 0.025 0.410 ± 0.057 0.410 ± 0.057 0.240 ± 0.052 0.300 ± 5.9e-17
Gemini 2.5 Pro 0.809 ± 0.023 0.856 ± 0.022 0.360 ± 0.117 0.420 ± 0.079 0.160 ± 0.070 0.360 ± 0.052
Gemini 3 Flash 0.866 ± 0.015 0.928 ± 0.015 0.300 ± 5.9e-17 0.400 ± 0.0 0.400 ± 0.0 0.500 ± 0.0
Gemini 3.1 Flash 0.844 ± 0.0 0.906 ± 0.0 0.400 ± 0.0 0.400 ± 0.0 0.200 ± 0.0 0.300 ± 5.9e-17
GPT-4.1 0.753 ± 0.023 0.806 ± 0.029 0.210 ± 0.088 0.260 ± 0.070 0.270 ± 0.048 0.340 ± 0.052
GPT-4o 0.856 ± 0.016 0.891 ± 0.022 0.130 ± 0.048 0.260 ± 0.084 0.0 ± 0.0 0.110 ± 0.032
GPT-5 0.787 ± 0.032 0.822 ± 0.033 0.350 ± 0.085 0.380 ± 0.063 0.230 ± 0.048 0.290 ± 0.032
GPT-5.2 0.784 ± 0.023 0.800 ± 0.026 0.280 ± 0.063 0.300 ± 0.067 0.370 ± 0.116 0.380 ± 0.132
GPT-5.2-Codex 0.731 ± 0.030 0.816 ± 0.023 0.320 ± 0.092 0.340 ± 0.084 0.280 ± 0.063 0.310 ± 0.032
GPT-5.3-Codex 0.753 ± 0.018 0.775 ± 0.025 0.200 ± 0.067 0.210 ± 0.074 0.190 ± 0.032 0.330 ± 0.048

ID Match (Rows) - Step-by-Step
Claude 3.7 0.794 ± 0.016 0.828 ± 0.030 0.230 ± 0.177 0.250 ± 0.165 0.150 ± 0.085 0.250 ± 0.085
Claude Opus 4.6 0.938 ± 0.0 0.969 ± 0.0 0.380 ± 0.042 0.440 ± 0.070 0.530 ± 0.048 0.590 ± 0.032
Claude Sonnet 4.5 0.906 ± 0.0 0.938 ± 0.0 0.400 ± 0.0 0.400 ± 0.0 0.200 ± 0.0 0.300 ± 5.9e-17
Gemini 2.5 Flash 0.853 ± 0.015 0.853 ± 0.015 0.200 ± 0.0 0.250 ± 0.053 0.190 ± 0.074 0.230 ± 0.048
Gemini 2.5 Pro 0.847 ± 0.023 0.875 ± 0.015 0.490 ± 0.032 0.500 ± 0.0 0.250 ± 0.053 0.370 ± 0.048
Gemini 3 Flash 0.938 ± 0.0 0.938 ± 0.0 0.390 ± 0.032 0.400 ± 0.0 0.300 ± 5.9e-17 0.400 ± 0.0
Gemini 3.1 Flash 0.812 ± 0.0 0.844 ± 0.0 0.300 ± 5.9e-17 0.300 ± 5.9e-17 0.100 ± 0.0 0.200 ± 0.0
GPT-4.1 0.797 ± 0.022 0.809 ± 0.034 0.270 ± 0.082 0.300 ± 0.115 0.210 ± 0.074 0.280 ± 0.079
GPT-4o 0.875 ± 0.033 0.900 ± 0.029 0.180 ± 0.042 0.270 ± 0.048 0.130 ± 0.067 0.220 ± 0.063
GPT-5 0.856 ± 0.030 0.884 ± 0.039 0.370 ± 0.048 0.390 ± 0.032 0.190 ± 0.032 0.260 ± 0.070
GPT-5.2 0.806 ± 0.029 0.809 ± 0.031 0.280 ± 0.042 0.290 ± 0.032 0.290 ± 0.074 0.290 ± 0.074
GPT-5.2-Codex 0.847 ± 0.045 0.869 ± 0.041 0.420 ± 0.092 0.440 ± 0.107 0.250 ± 0.053 0.280 ± 0.079
GPT-5.3-Codex 0.838 ± 0.013 0.875 ± 0.021 0.310 ± 0.074 0.320 ± 0.063 0.350 ± 0.053 0.360 ± 0.070

Column Match - Direct
Claude 3.7 0.738 ± 0.030 0.762 ± 0.034 0.680 ± 0.042 0.790 ± 0.057 0.310 ± 0.088 0.280 ± 0.092
Claude Opus 4.6 0.797 ± 0.016 0.812 ± 0.0 0.640 ± 0.052 0.790 ± 0.032 0.220 ± 0.042 0.220 ± 0.042
Claude Sonnet 4.5 0.781 ± 0.0 0.781 ± 0.0 0.500 ± 0.0 0.700 ± 0.0 0.300 ± 5.9e-17 0.300 ± 5.9e-17
Gemini 2.5 Flash 0.775 ± 0.025 0.800 ± 0.022 0.760 ± 0.052 0.820 ± 0.079 0.240 ± 0.052 0.350 ± 0.071
Gemini 2.5 Pro 0.828 ± 0.016 0.828 ± 0.016 0.760 ± 0.097 0.820 ± 0.063 0.200 ± 0.067 0.280 ± 0.042
Gemini 3 Flash 0.803 ± 0.015 0.812 ± 0.0 0.800 ± 0.0 0.800 ± 0.0 0.400 ± 0.0 0.500 ± 0.0
Gemini 3.1 Flash 0.812 ± 0.0 0.812 ± 0.0 0.900 ± 0.0 0.900 ± 0.0 0.300 ± 5.9e-17 0.300 ± 5.9e-17
GPT-4.1 0.709 ± 0.015 0.741 ± 0.015 0.600 ± 0.047 0.650 ± 0.053 0.230 ± 0.048 0.230 ± 0.048
GPT-4o 0.728 ± 0.026 0.738 ± 0.016 0.270 ± 0.082 0.450 ± 0.085 0.100 ± 0.0 0.120 ± 0.042
GPT-5 0.750 ± 0.026 0.775 ± 0.025 0.690 ± 0.110 0.760 ± 0.107 0.250 ± 0.071 0.280 ± 0.063
GPT-5.2 0.762 ± 0.034 0.778 ± 0.031 0.750 ± 0.071 0.770 ± 0.048 0.270 ± 0.067 0.300 ± 0.067
GPT-5.2-Codex 0.728 ± 0.033 0.809 ± 0.031 0.760 ± 0.084 0.790 ± 0.088 0.280 ± 0.042 0.320 ± 0.042
GPT-5.3-Codex 0.759 ± 0.030 0.797 ± 0.022 0.560 ± 0.052 0.560 ± 0.052 0.190 ± 0.057 0.330 ± 0.048

Column Match - Step-by-Step
Claude 3.7 0.734 ± 0.034 0.778 ± 0.027 0.530 ± 0.149 0.660 ± 0.143 0.280 ± 0.103 0.260 ± 0.052
Claude Opus 4.6 0.938 ± 0.0 0.938 ± 0.0 0.650 ± 0.053 0.720 ± 0.042 0.300 ± 5.9e-17 0.310 ± 0.032
Claude Sonnet 4.5 0.875 ± 0.0 0.875 ± 0.0 0.800 ± 0.0 0.900 ± 0.0 0.250 ± 0.053 0.300 ± 5.9e-17
Gemini 2.5 Flash 0.906 ± 0.0 0.906 ± 0.0 0.850 ± 0.053 0.880 ± 0.042 0.200 ± 0.0 0.240 ± 0.052
Gemini 2.5 Pro 0.906 ± 0.0 0.906 ± 0.0 0.790 ± 0.032 0.800 ± 0.047 0.240 ± 0.052 0.310 ± 0.057
Gemini 3 Flash 0.906 ± 0.0 0.906 ± 0.0 0.800 ± 0.0 0.810 ± 0.032 0.300 ± 5.9e-17 0.300 ± 5.9e-17
Gemini 3.1 Flash 0.875 ± 0.0 0.906 ± 0.0 0.800 ± 0.0 0.800 ± 0.0 0.300 ± 5.9e-17 0.300 ± 5.9e-17
GPT-4.1 0.903 ± 0.010 0.903 ± 0.010 0.680 ± 0.063 0.750 ± 0.053 0.250 ± 0.053 0.290 ± 0.032
GPT-4o 0.822 ± 0.021 0.822 ± 0.021 0.540 ± 0.097 0.700 ± 0.094 0.360 ± 0.070 0.390 ± 0.074
GPT-5 0.856 ± 0.037 0.872 ± 0.031 0.750 ± 0.097 0.830 ± 0.067 0.050 ± 0.053 0.180 ± 0.092
GPT-5.2 0.863 ± 0.016 0.863 ± 0.016 0.760 ± 0.052 0.770 ± 0.067 0.390 ± 0.074 0.390 ± 0.074
GPT-5.2-Codex 0.887 ± 0.037 0.900 ± 0.029 0.810 ± 0.057 0.830 ± 0.048 0.330 ± 0.067 0.380 ± 0.042
GPT-5.3-Codex 0.875 ± 0.0 0.906 ± 0.0 0.690 ± 0.057 0.730 ± 0.067 0.370 ± 0.067 0.380 ± 0.079

Notes. In each of the four sections of the table, the best mean per column is shown in boldface type, together with every LLM model whose run-level scores are not significantly worse than the best under a two-sided unpaired permutation test (Holm-corrected, α = 0.05).

Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.

Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.

Initial download of the metrics may take a while.