Model Releases
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
arXiv:2604.09029v1 Announce Type: cross Abstract: Large language models have been widely explored as decision-support tools in high-stakes domains due to their contextual understanding and reasoning c
arXiv:2604.09029v1 Announce Type: cross Abstract: Large language models have been widely explored as decision-support tools in high-stakes domains due to their contextual understanding and reasoning capabilities. However, existing decision-making benchmarks rely on two simplifying assumptions: actions are selected from a finite set of pre-defined candidates, and explicit conditions restricting action feasibility are not incorporated into the decision-making process. These assumptions fail to capture the compositional structure of real-world actions and the explicit conditions that constrain their validity. To address these limitations, we introduce CONDESION-BENCH, a benchmark designed to evaluate conditional decision-making in compositional action space. In CONDESION-BENCH, actions are defined as allocations to decision variables and are restricted by explicit conditions at the variable, contextual, and allocation levels. By employing oracle-based evaluation of both decision quality and condition adherence, we provide a more rigorous assessment of LLMs as decision-support tools.
Related
- Robust Reasoning Benchmark
- Medical Reasoning with Large Language Models: A Survey and MR-Bench
- Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization
- ToxReason: A Benchmark for Mechanistic Chemical Toxicity Reasoning via Adverse Outcome Pathway
- MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts
Source: arXiv cs.AI | 2026-04-13