TLA^{+}-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation
DGX agentarXiv:2607.23425v1 Announce Type: cross Abstract: Large language models increasingly write TLA^{+} formal specifications from natural-language descriptions, but progress is hard to measure: existing r