TorchLean: Formalizing Neural Networks in Lean
arXiv:2602.22631v2 Announce Type: replace-cross Abstract: Neural networks are increasingly deployed in scientific, safety critical, and mission critical pipelines, yet verification and analysis are of
Knowledge catalogue
arXiv:2602.22631v2 Announce Type: replace-cross Abstract: Neural networks are increasingly deployed in scientific, safety critical, and mission critical pipelines, yet verification and analysis are of
arXiv:2605.23929v1 Announce Type: new Abstract: Modern AI systems increasingly rely on workflows composed of multiple interacting agents, some powered by large language models (LLMs) and others by con
arXiv:2509.04445v2 Announce Type: replace Abstract: Recent AI trends seek to align AI models to learned human-centric objectives, such as personal preferences, utility, or societal values. Using stand
arXiv:2509.12672v2 Announce Type: replace Abstract: The volume of machine-generated content online has grown dramatically due to the widespread use of Large Language Models (LLMs), leading to new chal
arXiv:2605.24643v1 Announce Type: new Abstract: This paper presents reinforcement learning (RL) policies for dynamic quadrupedal locomotion in planetary exploration scenarios. Building on a taskoptimi
arXiv:2605.25612v1 Announce Type: cross Abstract: The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computat
arXiv:2605.23989v1 Announce Type: new Abstract: Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks
arXiv:2603.09581v2 Announce Type: replace Abstract: Adam is a widely used optimization algorithm in deep learning, yet the specific class of objective functions where it exhibits inherent advantages r
arXiv:2605.25731v1 Announce Type: new Abstract: Multi-trait essay scoring aims to provide fine-grained evaluation of writing quality across multiple dimensions. However, how to effectively post-train
arXiv:2605.25119v1 Announce Type: cross Abstract: Domain adaptation aims to mitigate performance degradation caused by distribution shifts between a labeled source domain and an unlabeled or sparsely
arXiv:2605.25110v1 Announce Type: cross Abstract: Aligning structured data is a fundamental problem in computer vision and machine learning, underlying tasks such as time series analysis, human action
arXiv:2605.24862v1 Announce Type: new Abstract: Cross-domain offline reinforcement learning (RL) aims to learn a policy in the target domain with a limited target domain dataset and a source domain da
arXiv:2603.29897v2 Announce Type: replace-cross Abstract: Reranking is a critical component in many information retrieval pipelines. Despite remarkable progress in text-only settings, multimodal reran
arXiv:2605.24977v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) often hallucinate findings when generating chest X-ray reports: they fabricate findings that are not present in
University of California STEM professors want standardized tests back due to severe math deficiencies among students: “We now observe preparation gaps so severe that instructors must reteach middle sc
Gary Marcus criticizes a company's prospectus as containing unserious, empty, and potentially dishonest claims before its inclusion in the S&P 500 index. The post appears to highlight concerns about i
arXiv:2605.25952v1 Announce Type: cross Abstract: Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable perfo
arXiv:2605.24449v1 Announce Type: cross Abstract: Although quadcopters boast impressive traversal capabilities enabled by their omnidirectional maneuverability, the need for continuous pilot control i
arXiv:2605.25517v1 Announce Type: new Abstract: AI answer engines generate answers from retrieved pages but cite only a few sources. This makes visibility depend not just on ranking, but on being cite
arXiv:2605.24202v1 Announce Type: new Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learn
arXiv:2605.25864v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewar
when you think about it SpaceX has cumulatively lost a lot less than OpenAI so maybe it’s a bargain? or maybe we just shouldn’t value companies at a trillion dollars until they actually show evidence
Visual Fold is a tool for simple visual organization of ComfyUI workflows that does not turn selected nodes into a subgraph or change workflow logic. Group folding and node alignment features enable c
arXiv:2602.06508v2 Announce Type: replace Abstract: Reinforcement learning (RL) can refine Vision-Language-Action (VLA) policies beyond behavior cloning, but real-world RL remains expensive due to ext
arXiv:2605.25044v1 Announce Type: new Abstract: Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-t
arXiv:2510.06672v3 Announce Type: replace Abstract: Reinforcement learning algorithms such as GRPO have driven recent advances in large language model (LLM) reasoning. While scaling the number of roll
100% agree. and this is important. and it will affect you, personally. The SpaceX IPO is the most brazen retail fleecing in modern market history. NASDAQ has REWRITTEN the index rules specifically for
arXiv:2605.05997v2 Announce Type: replace Abstract: Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vis
New York Times: A look at the UK's AI Safety Institute, whose researchers probe AI models for safety gaps, as its work becomes a blueprint for other governments' AI policies — The government's A.I. Se
arXiv:2605.23775v1 Announce Type: new Abstract: This study tackles the challenge of precise wood log counting, where applications of the proposed methodology can span from automated approaches for mat
arXiv:2602.05472v2 Announce Type: replace Abstract: The quest for expert-level reasoning in Large Language Models (LLMs) has been hampered by a persistent extit{reward bottleneck}: traditional reinfor
arXiv:2605.23562v1 Announce Type: cross Abstract: Sparse rewards are a major bottleneck in multi-agent reinforcement learning (MARL), where simultaneous learning induces non-stationarity and makes rew
As one of the first people to warn about a possible AI backlash—years ago—let me tell you this: it’s going to get much, much worse. It breaks my heart that AI—something I spent my whole life thinking
arXiv:2605.23234v1 Announce Type: new Abstract: Assessing the spatial fairness of predictive models involves establishing whether they are statistically penalizing (favoring) individuals associated wi
arXiv:2605.23500v1 Announce Type: new Abstract: Segmentation is a fundamental task in computer vision, underpinning pixel-level scene understanding and serving as a cornerstone for applications rangin
arXiv:2602.20102v2 Announce Type: replace-cross Abstract: Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe co
arXiv:2605.23780v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) need efficient mechanisms to update knowledge without degrading existing capabilities. While intrinsic multimod
arXiv:2602.11146v2 Announce Type: replace-cross Abstract: Preference optimization for diffusion and flow-matching models relies on reward functions that are both discriminatively robust and computatio
arXiv:2602.13985v2 Announce Type: replace Abstract: Artificial intelligence (AI) has demonstrated strong potential in clinical diagnostics, often achieving accuracy comparable to or exceeding that of
arXiv:2603.06610v2 Announce Type: replace Abstract: Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfor
arXiv:2512.11551v2 Announce Type: replace Abstract: The growing number of road users has significantly increased the risk of accidents in recent years. Vulnerable Road Users (VRUs) are particularly at
arXiv:2605.23471v1 Announce Type: cross Abstract: Aggressive driving is a major cause of traffic accidents and poses a serious threat to road safety. Although deep learning methods have shown promisin
arXiv:2605.23270v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consiste
arXiv:2605.23453v1 Announce Type: new Abstract: We conducted a reproducibility-oriented re-evaluation of prior migraine classification studies, correcting for data leakage and metric bias. We then int
arXiv:2605.23138v1 Announce Type: cross Abstract: Variational Quantum Algorithms (VQAs) potentially offer a pathway to practical quantum advantage, but their optimization is heavily hindered by barren
arXiv:2605.23326v1 Announce Type: new Abstract: We present ClimateChat-300K, a large-scale dataset of 299,329 public Facebook posts about climate change collected between May 2020 and May 2024 through
arXiv:2605.23264v1 Announce Type: cross Abstract: Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spectral misalig
arXiv:2605.23178v1 Announce Type: new Abstract: Despite recent progress, text-to-image models still struggle to generate semantically diverse and compositionally accurate multi-person interaction scen
arXiv:2605.23346v1 Announce Type: new Abstract: Discrete diffusion models have emerged as powerful frameworks for generating structured categorical data. However, efficiently sampling from reward-tilt
arXiv:2510.09136v2 Announce Type: replace-cross Abstract: Personalized news recommendations have become a standard feature of large news aggregation services, optimizing user engagement through automa
arXiv:2605.23315v1 Announce Type: cross Abstract: Large language models trained under diverse objectives and architectures have been shown to develop increasingly similar internal representations, an
arXiv:2605.23116v1 Announce Type: cross Abstract: Existing Video Anomaly Detection (VAD) methods typically rely on task-specific training, leading to strong domain dependency and high training costs.
👇@davidSacks raises some interesting points here, which deserve an answer, but arguably gets the solution exactly backwards. Are we not now already leaving unelected private companies precisely the ab
arXiv:2605.23797v1 Announce Type: cross Abstract: Aiming at identifying unexpected inputs from unknown classes, out-of-distribution (OOD) detection has emerged as a pivotal approach to enhancing the r
arXiv:2602.17653v2 Announce Type: replace Abstract: Recent work has shown that language models (LMs) trained on synthetic corpora can exhibit typological preferences that resemble cross-linguistic reg
arXiv:2502.17119v2 Announce Type: replace-cross Abstract: Deep generative models have made rapid progress in image, text, audio, and video generation, and are increasingly being applied to structured
arXiv:2605.23762v1 Announce Type: new Abstract: Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating
arXiv:2509.06858v2 Announce Type: replace-cross Abstract: Large Language Models are increasingly used to simulate human opinion dynamics, yet the effect of genuine interaction is often obscured by sys
arXiv:2605.23089v1 Announce Type: cross Abstract: Model-based reinforcement learning improves sample efficiency by learning a world model. However, existing latent world models such as DreamerV3 do no
arXiv:2605.23386v1 Announce Type: new Abstract: Agricultural UAV research requires simulators that integrate realistic 3D scenes, high-fidelity vehicle dynamics, and robotics middleware, while remaini