Jun Luo

h-index20
2papers
1,330citations

2 Papers

11.3AIJul 16
Can We Trust Item Response Theory for AI Evaluation?

Han Jiang, Sunbeom Kwon, Jinwen Luo et al.

AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities, rank systems, select informative examples, and diagnose benchmark quality. However, AI benchmark data often departs from the data regime of human testing, for which standard IRT estimation tools were originally developed: benchmarks typically involve fewer evaluated models, far more items, and capability distributions that may be skewed, clustered, or multimodal. We examine how these regime mismatches challenge the reliability of IRT modeling for AI evaluation. Using item parameters and capability distributions derived from six widely used LLM benchmarks, we simulate response matrices under three common IRT models and compare four estimation tools used in recent benchmark studies: marginal maximum likelihood, Markov chain Monte Carlo, variational inference, and a neural pseudo-Siamese estimator. Across 18,000 simulation conditions, we systematically evaluate computational feasibility, scalability, and the reliability of IRT inferences about model rankings, predicted performance, and item characteristics. Results show that classical estimators can become infeasible in large benchmark settings, whereas scalable estimators can produce unreliable item-level and ranking inferences with small or nonnormally distributed model sets. This study identifies when latent trait models reliably support or risk distorting AI benchmarking claims, and what sample sizes and diagnostics are needed for trustworthy use.

9.4QUANT-PHJul 14
Clifford-Only Quantum Reed-Solomon Codes and a Tornado Concatenation for Biased-Noise Cat Qubits

Cheng-You Ho, Justin Luo, Henry Ng et al.

Dissipative cat qubits exponentially suppress one Pauli error channel with the mean photon number, leaving the conjugate bit-flip error as the dominant failure mode. This strong noise bias makes the full machinery of general quantum error correction unnecessary: a code need only protect against a single error type, and any classical linear code can be promoted to a Clifford stabilizer code that does exactly this. We use this observation to build a Clifford-only quantum Reed-Solomon (RS) code. Starting from the [7,3,5] RS code over $GF(2^{3})$ which is maximum distance separable, we expand each field symbol into three bits to obtain the [21,9,6] linear code over $GF(2)$, realized as a [[21,9, $d_{X}=6$, $d_{Z}=1$]] bit-flip code whose stabilizers are products of Z operators. Because no phase-flip correction is attempted, the construction avoids the non-Clifford quantum Fourier transform required by the Grassl-Beth quantum RS codes and is fully simulable in Stim. Errors are decoded by a lookup table of minimum-weight corrections. We then introduce a Tornado architecture: a two-layer concatenation that wraps every position of the outer RS code in an inner distance-three repetition code, yielding a [[63, 9, 18]] code decoded in two stages, a majority vote within each repetition block followed by the outer lookup table. Monte Carlo simulations show that at a physical bit-flip rate $p=0.1$ the Tornado code reaches a logical error rate $p_{L}\approx5.3\times10^{-3}$, below both parent codes, and that its logical error rate scales as $p_{L}\propto p^{6}$ at low p, in contrast to $p^{2}$ for the repetition code and $p^{3}$ for the standalone RS code. We give the exact construction, the error and circuit model, an asymptotic scaling analysis, and an account of the overhead cost and of the assumptions behind the noise model.