Codec-Robust Attacks on Audio LLMs
arXiv:2605.20519v1 Announce Type: cross Abstract: Prior attacks on Audio Large Language Models (Audio LLMs) demonstrated that carefully crafted waveform-domain perturbations can force targeted adversa
Knowledge catalogue
arXiv:2605.20519v1 Announce Type: cross Abstract: Prior attacks on Audio Large Language Models (Audio LLMs) demonstrated that carefully crafted waveform-domain perturbations can force targeted adversa
arXiv:2605.22447v1 Announce Type: new Abstract: The study of online discourse has become central to understanding societal polarization. While much research has focused on detecting overt toxicity, th
arXiv:2605.22249v1 Announce Type: new Abstract: Accurate brain tumor segmentation using multiparametric MRI is critical for effective treatment planning. However, in clinical settings, complete acquis
arXiv:2605.21970v1 Announce Type: cross Abstract: Accurate and robust medical image classification is paramount for early disease diagnosis and treatment planning. However, challenges such as limited
arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is
arXiv:2605.21849v1 Announce Type: cross Abstract: Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers s
ha ha, so much for step change? maybe this problem was just easier than some? The standard GPT-5.5 reproduced the proof ~ 👇 https://chatgpt.com/share/6a0e9e04-8cb0-8332-a4f1-ec68acd2e03e You don't nee
Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a half ago. It has been lead by @reeselevine and team at USCS
arXiv:2605.22035v1 Announce Type: cross Abstract: Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge
arXiv:2510.00948v2 Announce Type: replace Abstract: Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent c
Link to GPT 5.5 on the recent Erdo problem: https://x.com/maxiao54704/status/2057484153755480537?s=61 The standard GPT-5.5 reproduced the proof ~ 👇 https://chatgpt.com/share/6a0e9e04-8cb0-8332-a4f1-ec
arXiv:2605.21954v1 Announce Type: new Abstract: Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal la
arXiv:2605.21796v1 Announce Type: cross Abstract: Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision
arXiv:2605.22550v1 Announce Type: new Abstract: Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags f
arXiv:2506.19500v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly act as function-call agents that invoke external tools to tackle tasks beyond their static knowledge
arXiv:2605.21401v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes
arXiv:2605.21776v1 Announce Type: new Abstract: Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether larg
arXiv:2511.08093v2 Announce Type: replace-cross Abstract: Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computa
arXiv:2605.22658v1 Announce Type: new Abstract: While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoni
arXiv:2509.17086v2 Announce Type: replace Abstract: Detecting and localizing poultry is essential for advancing smart poultry farming. Despite the progress of detection-centric methods, challenges per
arXiv:2605.21965v1 Announce Type: new Abstract: Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop
arXiv:2605.22158v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual token
arXiv:2605.22456v1 Announce Type: new Abstract: Cloud-hosted LLM driver agents provide useful semantic judgments, but their inference latency exceeds stepwise vehicle-control windows. Learned world mo
arXiv:2605.21198v1 Announce Type: cross Abstract: Public events on social media generate large volumes of discussion whose collective dynamics carry direct value for opinion forecasting and crisis res
A discussion of progressive training approaches for diffusion models where training begins at lower scale factors and progressively increases to target scale factors, leveraging previously trained mod
arXiv:2601.10348v2 Announce Type: replace Abstract: Efficient distillation is a key pathway for converting expensive reasoning capability into deployable efficiency, yet in the frontier regime where t
arXiv:2605.22355v1 Announce Type: new Abstract: Public transit route planning traditionally depends on structured map infrastructure and complex routing engines, and no existing dataset supports train
arXiv:2605.22620v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of LLMs, but often depends on external supervis
arXiv:2605.21611v1 Announce Type: new Abstract: We introduce spatially grounded contextual image generation, a controllable image generation task that reframes the conditioning paradigm. Instead of su
Another win for open-source robotics! 🔥 @huggingface just released a fully open-source humanoid robot, and you can build one for $2,500. I'm a huge advocate of open-source in robotics space. Why? Robo
arXiv:2605.21240v1 Announce Type: new Abstract: LLM agents have shown strong performance across a wide range of complex tasks, including interactive environments that require long-horizon decision mak
arXiv:2605.21063v1 Announce Type: new Abstract: Typical LLM responses tend to follow a default style, even though users often have distinct preferences regarding tone, verbosity, and formality that th
arXiv:2503.00565v3 Announce Type: replace-cross Abstract: The multi-armed bandits (MAB) framework is a widely used approach for sequential decision-making, where a decision-maker selects an arm in eac
arXiv:2605.20610v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models are often interpreted by analysing which categories are routed to which experts. However, routing alone does not reveal
arXiv:2509.07120v2 Announce Type: replace Abstract: Efficient and accurate feed-forward multi-view reconstruction has long been an important task in computer vision. Recent transformer-based models li
arXiv:2605.20975v1 Announce Type: new Abstract: Federated Learning enables collaborative model training across decentralized data sources without data transfer. Averaging-based FL is limited by the pr
arXiv:2605.21177v1 Announce Type: cross Abstract: This work presents extsc{ChunkFT}, a memory-efficient fine-tuning framework that reformulates full-parameter fine-tuning around a dynamically activate
arXiv:2605.20473v1 Announce Type: cross Abstract: Test-time scaling has emerged as a promising approach for improving code generation by exploring large solution spaces at inference time. However, exi
arXiv:2605.21049v1 Announce Type: new Abstract: Large language models (LLMs) reliably predict neural activity during language comprehension and transformer depth has been interpreted as mirroring hier
arXiv:2605.20771v1 Announce Type: new Abstract: Spurious correlations in real-world datasets cause machine learning models to rely on irrelevant patterns, undermining reliability, generalization, and
arXiv:2605.20576v1 Announce Type: new Abstract: Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing
arXiv:2605.20982v1 Announce Type: cross Abstract: AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of mitigations
arXiv:2605.20743v1 Announce Type: cross Abstract: Vision-language models solve geometry problems with rising accuracy, yet their intermediate states remain latent and unverifiable: a relation expresse
arXiv:2605.21028v1 Announce Type: new Abstract: Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term continuity w
arXiv:2605.20254v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research attention, bec
arXiv:2605.20787v1 Announce Type: new Abstract: The rapid advancements in generative AI technologies, such as Stable Diffusion, DALL-E, and Midjourney, have significantly transformed the creation of s
arXiv:2605.20744v1 Announce Type: new Abstract: Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby
arXiv:2605.21309v1 Announce Type: new Abstract: Cooperative perception enabled by Vehicle-to-Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental repre
arXiv:2605.21363v1 Announce Type: new Abstract: As large language models (LLMs) increasingly shape how users form, refine, and extend their goals, attributing contributions in human-AI collaboration b
arXiv:2605.20257v1 Announce Type: new Abstract: Recently, instance discrimination models have emerged as a major solution for self-supervised learning. Having already demonstrated its effectiveness in
arXiv:2605.20244v1 Announce Type: cross Abstract: We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refactoring of L
arXiv:2605.20667v1 Announce Type: new Abstract: Detecting small unmanned aerial vehicles from RGB-infrared remote-sensing pairs remains challenging due to tiny target scale, cluttered backgrounds, and
📣Meet Qwen3.7-Max — our latest flagship, made for the Agent Era. A versatile foundation for agents that actually get things done: 🧑💻 Coding agent, end to end. Frontend prototypes, multi-file refactor
arXiv:2605.21463v1 Announce Type: new Abstract: We present Mem-pi, a framework for adaptive memory in large language model (LLM) agents, where useful guidance is generated on demand rather than retrie
arXiv:2605.20833v1 Announce Type: new Abstract: Memory is a central capability for LLM agents operating across long-horizon tasks. Existing memory benchmarks predominantly evaluate retention of person
🦔Microsoft canceled its internal Claude Code licenses this week after token-based billing made the cost untenable, even for a company with effectively infinite cloud resources. Uber's CTO sent an inte
arXiv:2602.24138v2 Announce Type: replace Abstract: Automated recognition of surgical phases and steps is a fundamental capability for intraoperative decision support, workflow automation, and skill a
arXiv:2605.21437v1 Announce Type: cross Abstract: Standard approaches to forecasting the weekly number of earthquakes on a spatial grid rely on the Poisson distribution with a single global dispersion
arXiv:2605.20525v1 Announce Type: cross Abstract: We present NeuroQA, a large-scale benchmark for visual question answering in 3D brain magnetic resonance imaging (MRI), with 56,953 QA pairs from 12,9
arXiv:2605.21322v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative model training without centralizing data. However, real-world deployments must simultaneously address stat