Learning How to Cube
arXiv:2605.16632v1 Announce Type: cross Abstract: Despite the effectiveness of Cube-and-Conquer (C&C) for solving challenging Boolean Satisfiability (SAT) problems, no prior work has shown that transf
Knowledge catalogue
arXiv:2605.16632v1 Announce Type: cross Abstract: Despite the effectiveness of Cube-and-Conquer (C&C) for solving challenging Boolean Satisfiability (SAT) problems, no prior work has shown that transf
arXiv:2605.17003v1 Announce Type: cross Abstract: Reinforcement Learning (RL) post-training has emerged as the dominant paradigm for eliciting mathematical reasoning in Large Language Models (LLMs), y
arXiv:2605.16474v1 Announce Type: cross Abstract: The integration of advertising auction mechanisms into large language model (LLM)-based chatbots presents a significant opportunity for commercializat
arXiv:2605.18541v1 Announce Type: new Abstract: Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling,
arXiv:2410.13846v3 Announce Type: replace-cross Abstract: Scaling language models to handle longer contexts introduces substantial memory challenges due to the growing cost of key-value (KV) caches. M
arXiv:2309.05646v2 Announce Type: replace-cross Abstract: Distributed Denial of Service (DDoS) attacks remain a persistent threat to the availability of Internet services, edge networks, and cyber-phy
arXiv:2605.16675v1 Announce Type: new Abstract: We introduce LinAlg-Bench, a diagnostic benchmark evaluating 10 frontier large language models on structured linear algebra computation across a strict
arXiv:2603.00631v2 Announce Type: replace Abstract: LiTS is a modular Python framework for LLM reasoning via tree search. It decomposes tree search into three reusable components (Policy, Transition,
Live from Code with Claude London: we're launching self-hosted sandboxes (public beta) and MCP tunnels (research preview) in Claude Managed Agents. Run agents inside your own perimeter, with your secu
arXiv:2605.17986v1 Announce Type: cross Abstract: AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injection (IPI) ri
llm-gemini 0.32 is an alpha release of Simon Willison's LLM Python library and CLI tool that provides access to Google's Gemini models , continuing work on major architectural changes to support newer
I don't have current information about this specific entry, so I'll describe what it likely covers based on the available details. This entry documents the release or update of llm-gemini version 0.32
arXiv:2605.17653v1 Announce Type: cross Abstract: Sub-billion-parameter Transformer language models are increasingly deployed on edge devices, where the privacy, latency, and operating-cost advantages
arXiv:2605.16538v1 Announce Type: cross Abstract: This paper examines the opportunities, limitations, and practical considerations associated with the use of large language models (LLMs) in qualitativ
arXiv:2605.18565v1 Announce Type: cross Abstract: Real-world agents operate over long and evolving horizons, where information is repeatedly updated and may interfere across memories, requiring accura
arXiv:2605.16343v1 Announce Type: cross Abstract: Looped language models (LoopLMs) improve parameter efficiency by recursively reusing Transformer blocks, enabling deeper computation under a fixed mod
arXiv:2605.16375v1 Announce Type: new Abstract: Accurate air quality prediction is essential for public health, environmental monitoring, and industrial safety. However, most existing approaches rely
arXiv:2605.18253v1 Announce Type: cross Abstract: Recent masked diffusion language models (MDLMs), such as LLaDA and Dream, have achieved performance comparable to autoregressive large language models
arXiv:2605.17159v1 Announce Type: new Abstract: Document processing automation remains a critical challenge in enterprise environments, where traditional manual approaches are labor-intensive and erro
arXiv:2605.18617v1 Announce Type: cross Abstract: Most existing vision-language manipulation research targets rigid robotic arms, whose fixed morphology limits adaptability in cluttered or confined sp
arXiv:2605.16301v1 Announce Type: cross Abstract: Single-turn benchmarks such as AnimalHarmBench (AHB) have established important baselines for measuring animal welfare alignment in large language mod
arXiv:2605.18176v1 Announce Type: cross Abstract: This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participa
arXiv:2605.16716v1 Announce Type: cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single pr
arXiv:2605.16290v1 Announce Type: cross Abstract: Predicting the difficulty of multiple-choice questions (MCQs) is important for effective assessment, yet current methods typically assume a unimodal s
arXiv:2603.05308v2 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification. While large language model
arXiv:2605.16445v1 Announce Type: cross Abstract: Masked Diffusion Language Models MDLMs replace autoregressive generation with iterative demasking and their privacy properties are largely unstudied.
arXiv:2601.21468v5 Announce Type: replace Abstract: Long-horizon agentic reasoning necessitates effectively compressing growing interaction histories into a limited context window. Most existing memor
arXiv:2602.12871v2 Announce Type: replace Abstract: Large language models (LLMs) have attracted growing interest as supportive tools for psychiatric assessment and clinical decision support. However,
arXiv:2510.10528v3 Announce Type: replace Abstract: Large reasoning models (LRMs) have demonstrated remarkable proficiency in tackling complex tasks through step-by-step thinking. However, this length
arXiv:2605.17292v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems have shown promise for solving complex tasks through agent collaboration. However, existing frameworks as
arXiv:2605.17398v1 Announce Type: new Abstract: This paper presents MiniGPT, a compact from-scratch implementation of GPT-style autoregressive language modeling in PyTorch. The aim is to rebuild the c
arXiv:2605.17198v1 Announce Type: cross Abstract: To be useful for downstream applications, vision decoding models that are trained to reconstruct seen images from human brain activity must be able to
arXiv:2601.08118v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning da
Reuters: Mistral acquires Vienna-based Emmi AI for an undisclosed sum to boost its industrial offerings in Europe; Emmi raised €15M in Austria's largest round in 2025 — Europe's leading artificial int
arXiv:2605.16865v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is widely used to inject new knowledge into language models, but it often degrades pretrained capabilities such as reasonin
arXiv:2506.12119v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token co
arXiv:2605.17598v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures enable efficient model scaling, yet expert routing behavior across underrepresented languages remains poorly unde
arXiv:2605.16616v1 Announce Type: new Abstract: Autonomous research systems capable of generating complete scientific manuscripts have advanced rapidly, yet robust and realistic evaluation frameworks
arXiv:2602.22667v2 Announce Type: replace Abstract: Open-vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abunda
arXiv:2511.17392v3 Announce Type: replace Abstract: Deformable image registration (DIR) remains a fundamental yet challenging problem in medical image analysis, largely due to the prohibitively high-d
arXiv:2605.17454v1 Announce Type: new Abstract: Multi-party multi-objective optimization problems (MPMOPs) require consensus among autonomous decision makers and therefore differ from flattened many-o
arXiv:2605.17859v1 Announce Type: cross Abstract: Wearables are widely used for mobile health monitoring, and photoplethysmography (PPG) is a key sensing modality for heart rate and related physiologi
arXiv:2605.18239v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attempts that circumvent safety guardrails. We investigate whether multi-turn conversation
arXiv:2605.16409v1 Announce Type: cross Abstract: Optical character recognition (OCR) and multilingual text understanding remain major failure modes of multimodal large language models (MLLMs), partic
arXiv:2605.17669v1 Announce Type: new Abstract: The preservation and interpretation of cultural heritage increasingly rely on digital technologies, among which Knowledge Graphs (KGs) stand out for the
Google's Gemini 3.5 Flash model costs approximately 3x more than Gemini 3 Flash, despite being a newer version. Google plans to integrate Gemini 3.5 Flash into many of their own products, suggesting t
arXiv:2605.16757v1 Announce Type: new Abstract: Multi-agent language systems are often built as hand-designed workflows, where agents are assigned semantic roles and communication protocols are specif
arXiv:2605.16923v1 Announce Type: new Abstract: Decoding visual information from electroencephalography (EEG) signals remains a fundamental challenge in brain-computer interfaces and medical rehabilit
arXiv:2605.17596v1 Announce Type: new Abstract: We present NeuSymMS, an adaptive memory system that enables large language model (LLM) agents to learn, remember, and reason about users across sessions
New upgrades to the @GeminiApp are you helping you get more done: ✨Gemini Spark is your 24/7 personal AI agent that can take action on your behalf, under your direction. It seamlessly integrates with
arXiv:2605.17364v1 Announce Type: new Abstract: Media bias detection has predominantly been framed as a classification task: assign a political label to an article or outlet. We argue this framing is
arXiv:2605.16317v1 Announce Type: new Abstract: Accurate, unified models for event cameras (ECs) remain elusive, hampering calibration and algorithm design. We develop a foundational probabilistic mod
arXiv:2605.16537v1 Announce Type: new Abstract: Open-source mobile manipulators have reached 660 (XLeRobot) but every sub-1,000 platform shares three limitations: a fixed-height workspace, reactive-on
arXiv:2605.18593v1 Announce Type: cross Abstract: Open-vocabulary embodied AI agents increasingly rely on vision-language models such as CLIP for object perception and task grounding. However, the sha
This blog post from Anthropic's Claude discusses the practical effectiveness of using HTML within Claude's code interpreter, likely exploring how HTML can be leveraged for web development, data visual
arXiv:2605.18509v1 Announce Type: new Abstract: Automated decision-making algorithms drive applications such as recommendation systems and search engines. These algorithms often rely on off-policy con
arXiv:2605.17360v1 Announce Type: new Abstract: Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming
arXiv:2602.02262v3 Announce Type: replace-cross Abstract: LLM-powered coding agents are redefining how real-world software is developed. To drive the research towards better coding agents, we require
arXiv:2605.18577v1 Announce Type: new Abstract: Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emer
arXiv:2605.16962v1 Announce Type: cross Abstract: Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the