ABRA: Agent Benchmark for Radiology Applications
arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo
Knowledge catalogue
arXiv:2605.11224v1 Announce Type: new Abstract: Existing medical-agent benchmarks deliver imaging as pre-selected samples, never as an environment the agent must navigate. We introduce ABRA, a radiolo
arXiv:2211.03524v2 Announce Type: replace Abstract: Modern Review Helpfulness Prediction systems are dependent upon multiple modalities, typically texts and images. Unfortunately, those contemporary a
arXiv:2605.10987v1 Announce Type: new Abstract: Modern machine learning deployments increasingly compose specialized models into dynamic inference pipelines, where upstream components produce intermed
arXiv:2605.12375v1 Announce Type: new Abstract: Accurate crop yield forecasting in commercial soft fruit production is constrained by the data available in typical commercial farm records, which lack
arXiv:2605.12316v1 Announce Type: new Abstract: We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification
arXiv:2501.16931v2 Announce Type: replace Abstract: Machine learning models are often evaluated using point estimates of performance metrics such as accuracy, F1 score, or mean squared error. Such sum
arXiv:2605.11829v1 Announce Type: cross Abstract: The accurate recovery of constituent-level optical properties from integrating sphere measurements is a central analytical challenge in pharmaceutical
arXiv:2605.11533v1 Announce Type: new Abstract: Clinical check-up reports are multimodal documents that combine page layouts, tables, numerical biomarkers, abnormality flags, imaging findings, and dom
arXiv:2605.11143v1 Announce Type: new Abstract: Reasoning benchmarks measure clinical performance on clean inputs. We evaluate the step before reasoning: retrieval over real EHR notes, where negation,
arXiv:2605.11485v1 Announce Type: new Abstract: Imitation learning powered by generative models has proven effective for modeling complex single-agent behaviors. However, teaching multi-agent systems,
arXiv:2605.11210v1 Announce Type: new Abstract: We present a framework for distributed Pose Graph Optimization (PGO) by formulating the problem as a second-order continuous-time dynamical system evolv
arXiv:2506.13163v3 Announce Type: replace Abstract: We study the Logistic Contextual Slate Bandit problem, where, at each round, an agent selects a slate of N items from an exponentially large set (of
arXiv:2504.14707v3 Announce Type: replace Abstract: We introduce FLAME (FLemish Accounts of Momentary Experiences), a new corpus of nearly 25,000 daily personal narratives in Belgian-Dutch (Flemish),
arXiv:2605.11706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to complete complex tasks by selecting and coordinating external tools across multiple steps. This re
arXiv:2605.11206v1 Announce Type: new Abstract: Instructions trigger a production-centered mechanism in language models. Through a cognitively inspired lens that separates language processing and prod
arXiv:2505.16156v3 Announce Type: replace-cross Abstract: Quantifying differences between probability distributions is fundamental to statistics and machine learning, primarily for comparing statistic
arXiv:2605.11385v1 Announce Type: new Abstract: Stochastic Human Trajectory Prediction (HTP) using generative modeling has emerged as a significant area of research. Although state-of-the-art models e
arXiv:2605.12156v1 Announce Type: new Abstract: Automatic misinformation detection performs well when deception is visible in what an article explicitly states. However, some misinformation articles r
arXiv:2605.11262v1 Announce Type: new Abstract: Chain-of-thought and more broadly test-time compute are known to augment the expressive capabilities of language models and have led to major innovation
arXiv:2602.05830v2 Announce Type: replace-cross Abstract: Floating-point neural networks dominate modern machine learning but incur substantial inference costs, motivating emerging interest in Boolean
arXiv:2605.12084v1 Announce Type: cross Abstract: Designing learnable information-theoretic objectives for robot exploration remains challenging. Such objectives aim to guide exploration toward data t
literally what i have been saying for years, once again. Yann LeCun says you cannot build a reliable agentic system without a world model LLMs don't have world models. They can't predict the consequen
arXiv:2602.09587v2 Announce Type: replace Abstract: The scarcity of high-quality data remains a primary bottleneck in adapting multimodal generative models for medical image editing. Existing medical
arXiv:2605.11471v1 Announce Type: new Abstract: Matrix product operator Born machines (MPO-BMs) are tractable tensor-network models for probabilistic modeling, but their efficient approximation capabi
Open Source always wins 👀 Save your credits and use open source models when you can. The new LTX 2.3 lipdub LoRA paired with Chatterbox TTS voice cloning model is the best workflow for lip syncing, ch
arXiv:2605.11652v1 Announce Type: cross Abstract: We study posterior contraction rates for sparse Bayesian Kolmogorov-Arnold networks (KANs) over anisotropic Besov spaces, providing a statistical foun
arXiv:2605.11534v1 Announce Type: new Abstract: When an LLM-based embodied agent fails at a household task, the culprit could be misidentified objects, forgotten sub-goals, or poor action sequencing -
arXiv:2512.24558v2 Announce Type: replace-cross Abstract: Neural quantum states efficiently represent many-body wavefunctions with neural networks, but the cost of Monte Carlo sampling limits their sc
arXiv:2504.12326v3 Announce Type: replace Abstract: Clinical case reports and discharge summaries may be the most complete and accurate summarization of patient encounters, yet they are finalized, i.e
arXiv:2605.11462v1 Announce Type: new Abstract: Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle
arXiv:2511.11935v2 Announce Type: replace Abstract: Deep-learning survival models for electronic health record (EHR) data are hard to compare across papers because the upstream preprocessing step, whi
arXiv:2605.12487v1 Announce Type: new Abstract: We explore the effectiveness of an LLM-guided query refinement paradigm for extending the usability of embedding models to challenging zero-shot search
arXiv:2605.05971v2 Announce Type: replace Abstract: Long-context language modeling is increasingly constrained by the Key-Value (KV) cache, whose memory and decode-time access costs scale linearly wit
NVIDIA's video analytics AI agents analyze and process large volumes of video data through natural language tasks to provide critical insights , powered by vision language models, large language model
arXiv:2605.11334v1 Announce Type: cross Abstract: LLM-as-Judge systems are widely deployed for automated evaluation, yet practitioners lack reliable methods to know when a judge's verdict should be tr
arXiv:2605.12112v1 Announce Type: new Abstract: RLHF is widely used to align flow-matching text-to-image models with human preferences, but often leads to severe diversity collapse after fine-tuning.
arXiv:2605.11240v1 Announce Type: cross Abstract: Generative AI models differ from traditional machine learning tools in that they allow users to provide as much or as little information as they choos
arXiv:2605.08600v1 Announce Type: new Abstract: We present a new publicly available corpus of 100,502 movie reviews from Kazakhstan collected from kino.kz, spanning 2001-2025 and covering 4,943 unique
arXiv:2605.09034v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization has become increasingly popular and important in fine-tuning large language models (LLMs), especially on edge devices due
arXiv:2510.18184v3 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse aut
arXiv:2505.18184v2 Announce Type: replace-cross Abstract: The increase in cardiac and pulmonary diseases presents an alarming and pervasive health challenge on a global scale responsible for unexpecte
arXiv:2603.29981v2 Announce Type: replace Abstract: Reliable estimation of predictive performance is essential for spatial environmental modeling, where machine-learning models are used to generate ma
arXiv:2605.09698v1 Announce Type: new Abstract: As data-science agents shift from co-pilots to auto-pilots, silent misframing becomes a critical failure mode. Agents quietly commit to plausible but un
arXiv:2605.08764v1 Announce Type: cross Abstract: Deep vision models degrade sharply in low-data regimes, particularly in medical imaging where labeled samples are scarce. We show this arises not mere
arXiv:2605.10397v1 Announce Type: cross Abstract: Visual anomaly detection (VAD) is crucial in many real-world fields, such as industrial inspection, medical imaging, infrastructure monitoring, and re
arXiv:2605.10480v1 Announce Type: new Abstract: Over the years, research in system identification has provided a rich set of methods for learning dynamical models, together with well-established theor
arXiv:2605.09533v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly employed in enterprise question-answering (QA) systems, requiring adaptation to domain-specific knowledg
arXiv:2605.10034v1 Announce Type: new Abstract: Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driv
arXiv:2605.08891v1 Announce Type: new Abstract: Sparse autoencoders have become a standard tool for uncovering interpretable latent representations in neural networks. Yet salient concepts often span
arXiv:2504.21228v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are susceptible to indirect prompt injection attacks, where the model inadvertently responds to instructions inje
arXiv:2601.12369v3 Announce Type: replace Abstract: Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing the
arXiv:2605.09092v1 Announce Type: new Abstract: This study addresses automatic transliteration from Tajik (Cyrillic script) to Persian (Perso-Arabic script). We present a curated, lexicographically ve
arXiv:2605.08399v1 Announce Type: new Abstract: Tool-augmented language models can extend small language models with external executable skills, but scaling the tool library creates a coupled challeng
ComfyUI support for HiDream-O1-Image enables local image generation with text prompts and optional reference images, featuring various precision options (BF16/FP16/FP32/FP8) and integration with atten
arXiv:2605.10673v1 Announce Type: new Abstract: Low-bit forward evaluation is an attractive route to memory-efficient zeroth-order (ZO) adaptation: the optimizer needs only scalar losses, and the mode
arXiv:2605.10787v1 Announce Type: new Abstract: Current LLM agents are proficient at calling isolated APIs but struggle with the 'last mile' of commercial software automation. In real-world scenarios,
arXiv:2605.09126v1 Announce Type: new Abstract: Asynchronous DiLoCo systems may receive pseudo-gradients computed several outer rounds earlier, yet the standard Nesterov outer optimizer does not expli
arXiv:2605.10887v1 Announce Type: new Abstract: Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far
arXiv:2605.09548v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress in mathematical reasoning, but this ability is not equally accessible across languages. E
arXiv:2605.08794v1 Announce Type: cross Abstract: Modern generative models can be understood as probability transport from a simple base distribution to a target data distribution. Deterministic trans