CompanyAnthropic2 recent entries13 May 2026ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IVarXiv:2605.11143v1 Announce Type: new Abstract: Reasoning benchmarks measure clinical performance on clean inputs. We evaluate the step before reasoning: retrieval over real EHR notes, where negation,→29 May 2026FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning BenchmarksarXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro
CompanyMeta1 recent entries9 Jun 2026AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel SynthesisarXiv:2606.09682v1 Announce Type: new Abstract: AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one
CompanyDeepSeek2 recent entries30 Apr 2026DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video AnnotationarXiv:2604.26565v1 Announce Type: new Abstract: Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora,→29 May 2026FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning BenchmarksarXiv:2605.29001v1 Announce Type: cross Abstract: A paraphrase-quality audit of MathCheck (ICLR 2025) detected 4 semantically incorrect paraphrases in 129 groups (3.1%); removing them drops GPT-4o fro
CompanyNVIDIA1 recent entries9 Jun 2026AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel SynthesisarXiv:2606.09682v1 Announce Type: new Abstract: AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one