AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation
arXiv:2607.00062v1 Announce Type: cross Abstract: High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason about algorith