Deep Double Q-learning
arXiv:2507.00275v2 Announce Type: replace-cross Abstract: Double Q-learning is a classical control algorithm that mitigates the maximization bias of Q-learning. To do so, it explicitly trains two inde
Knowledge catalogue
arXiv:2507.00275v2 Announce Type: replace-cross Abstract: Double Q-learning is a classical control algorithm that mitigates the maximization bias of Q-learning. To do so, it explicitly trains two inde
arXiv:2605.15202v1 Announce Type: new Abstract: Presentations are a primary medium for scholarly communication, yet most AI slide generators optimize the artifact (a visually plausible deck) while und
arXiv:2605.15532v1 Announce Type: cross Abstract: Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process are typically
arXiv:2605.16255v1 Announce Type: cross Abstract: Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major chall
arXiv:2605.15569v1 Announce Type: cross Abstract: Microservices are widely adopted in modern cloud systems due to their scalability and fault tolerance. However, microservice architectures introduce s
arXiv:2605.15967v1 Announce Type: new Abstract: We study event-graph substrates: a class of world models that represent agent state as an append-only log of typed RDF triples and answer counterfactual
arXiv:2605.15410v1 Announce Type: cross Abstract: Adaptive Non-local Observables (ANOs) have shown that making quantum observables dynamic can substantially enlarge the function space of Variational Q
arXiv:2605.15460v1 Announce Type: cross Abstract: Cross-modal hashing enables efficient retrieval by encoding images and text into compact binary codes. State-of-the-art methods rely on semantic simil
arXiv:2605.15519v1 Announce Type: cross Abstract: Visual active search (VAS) has been introduced as a modeling framework that leverages visual cues to direct aerial (e.g., UAV-based) exploration and p
arXiv:2605.15725v1 Announce Type: cross Abstract: Latent Action Models (LAMs) enable the learning of world models from unlabeled video by inferring abstract actions between consecutive frames. However
arXiv:2605.15225v1 Announce Type: cross Abstract: Biologically-inspired AI agent frameworks claim reliability benefits through structural guarantees adapted from gene regulatory networks, immune syste
arXiv:2504.00289v3 Announce Type: replace-cross Abstract: The release of top-performing open-weight LLMs has cemented China's role as a leading force in AI development. Do these models support languag
arXiv:2605.15205v1 Announce Type: new Abstract: Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and h
arXiv:2605.15543v1 Announce Type: cross Abstract: Many games of interest in the real world are often intractably large, thereby necessitating the use of game abstraction to shrink them in size, typica
arXiv:2511.19399v3 Announce Type: replace-cross Abstract: Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are tr
arXiv:2512.09673v3 Announce Type: replace-cross Abstract: Equivariant neural networks encode the intrinsic symmetry of data as an inductive bias, which has achieved impressive performance in wide doma
arXiv:2605.15542v1 Announce Type: new Abstract: GUI agents powered by Multimodal Large Language Models (MLLMs) have demonstrated impressive capability in understanding and executing user instructions.
arXiv:2605.15461v1 Announce Type: cross Abstract: Building state-of-the-art (SOTA) predictive models for drug discovery requires expensive search over tools, architectures, and training strategies. Cu
arXiv:2509.23352v3 Announce Type: replace-cross Abstract: The integration of Reinforcement Learning (RL) into flow matching models for text-to-image (T2I) generation has driven substantial advances in
arXiv:2605.15221v1 Announce Type: cross Abstract: AlphaEvolve and FunSearch have demonstrated the potential of combining large language models (LLMs) with evolutionary search for automated algorithm d
arXiv:2605.15586v1 Announce Type: cross Abstract: Complementary-label learning (CLL) is a weakly supervised paradigm where instances are labeled with classes they do not belong to. Despite a decade of
arXiv:2512.15067v3 Announce Type: replace-cross Abstract: The rapid growth in wireless infrastructure has increased the need to accurately estimate and forecast electromagnetic field (EMF) levels to e
arXiv:2605.15377v1 Announce Type: new Abstract: As AI systems are increasingly deployed in autonomous agentic settings at scale, it is important to ensure the actions they take are safe and aligned wi
arXiv:2605.12581v1 Announce Type: cross Abstract: Synthesising autonomous agents that can navigate uncertain environments while adhering to complex temporal constraints remains a fundamental challenge
arXiv:2605.16126v1 Announce Type: cross Abstract: For a fixed flow-based generative model under a small inference budget, sample quality can depend strongly on where the sampler spends its few functio
arXiv:2605.16223v1 Announce Type: cross Abstract: Generative video models are increasingly used in design animation tasks, yet no standardized evaluation framework exists for this domain. Unlike natur
arXiv:2601.23068v2 Announce Type: replace-cross Abstract: Computing the importance of features in supervised classification tasks is critical for model interpretability. Shapley values are a widely us
arXiv:2605.15417v1 Announce Type: cross Abstract: In GFlowNets and variational inference, it has been shown that the mean square error between target and model log probabilities is an effective, low v
arXiv:2605.15217v1 Announce Type: new Abstract: Instruction-tuned language models exhibit behavioural fairness in high-stakes decisions while retaining biased associations in their internal representa
arXiv:2505.21535v4 Announce Type: replace-cross Abstract: While transformers dominate modern vision and language models, their attention mechanism remains poorly suited for in-memory computing (IMC) d
arXiv:2605.15212v1 Announce Type: cross Abstract: We propose a new numerical method to estimate the fault tolerance of failure modes in digital circuit structures with a generative network sampling te
arXiv:2605.16099v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative training across decentralized clients, but most methods assume aligned feature schemas, an assumption th
arXiv:2605.15705v1 Announce Type: cross Abstract: World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become un
arXiv:2410.02832v2 Announce Type: replace-cross Abstract: This paper proposes a simple yet effective jailbreak attack named FlipAttack against black-box LLMs. First, from the autoregressive nature, we
arXiv:2602.23566v2 Announce Type: replace-cross Abstract: We study generative modeling of graphs with recurring subgraph motifs. We propose Flowette, a continuous flow matching framework that employs
arXiv:2312.05975v3 Announce Type: replace-cross Abstract: Explainability is a vital aspect of modern AI for real-world impact and usability. The main objective of this paper is to emphasise the need t
arXiv:2605.16233v1 Announce Type: new Abstract: Can LLM agents improve decision-making through self-generated memory without gradient updates? We propose FORGE (Failure-Optimized Reflective Graduation
arXiv:2605.16198v1 Announce Type: new Abstract: We examine one particular dimension of AI governance: how to monitor and audit AI-enabled products and services throughout the AI development lifecycle,
arXiv:2603.16011v2 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to op
arXiv:2605.15299v1 Announce Type: cross Abstract: In search and recommendation systems, predictive models often suffer from temporal instability when certain input features introduce volatility in out
arXiv:2605.15412v1 Announce Type: cross Abstract: Modern quantitative trading increasingly relies on systematic models to extract predictive signals from large-scale financial data, where alpha factor
arXiv:2605.16026v1 Announce Type: cross Abstract: Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performan
arXiv:2508.20810v3 Announce Type: replace Abstract: Rigorous evaluation of domain-specific language models requires benchmarks that are comprehensive, contamination-resistant, and maintainable. Static
arXiv:2605.15334v1 Announce Type: cross Abstract: The automatic synthesis of a program from any form of specification is regarded as a holy grail of computer science. Fueled by LLMs, NL2Code has achie
arXiv:2605.15445v1 Announce Type: new Abstract: Automated proving of polynomial inequalities is a fundamental challenge in automated mathematical reasoning, where rich algebraic structure and a rapidl
arXiv:2511.09378v2 Announce Type: replace Abstract: A series of influential studies established that large language models cannot reliably solve even simple planning tasks. We show that the latest gen
arXiv:2605.15880v1 Announce Type: cross Abstract: Thermal infrared imaging is robust to illumination variations and smoke interference, making it important for all-weather perception. However, the lac
arXiv:2605.16215v1 Announce Type: new Abstract: Clinical decision support systems (CDSS) require scrutable, auditable pipelines that enable rigorous, reproducible validation. Yet current LLM-based CDS
arXiv:2605.15836v1 Announce Type: cross Abstract: Learning visuomotor policies from scarce expert demonstrations remains a core challenge in robotic manipulation. A primary hurdle lies in distilling h
arXiv:2605.15223v1 Announce Type: cross Abstract: This paper presents an LLM-empowered workflow for RISC-V supply chain analysis, integrating Vision-Language Models (VLMs) and Model-Driven Engineering
arXiv:2508.17218v3 Announce Type: replace-cross Abstract: Reliable path planning in stochastic transportation networks requires decisions that account for uncertain and correlated travel times on irre
arXiv:2605.15905v1 Announce Type: cross Abstract: Modeling long-term user interests with massive historical user behaviors enhances click-through rate (CTR) prediction performance in advertising and r
arXiv:2605.16122v1 Announce Type: cross Abstract: Diffusion-based image synthesis has made AI-generated images (AIGI) increasingly photorealistic, raising urgent concerns about authenticity in applica
arXiv:2605.16094v1 Announce Type: cross Abstract: Wideband channel estimation (CE) in high-mobility scenarios remains challenging because channel responses vary rapidly, while practical systems can al
arXiv:2605.15295v1 Announce Type: cross Abstract: Machine learning (ML) algorithms are increasingly deployed in high-stakes decision-making domains such as loan approvals, hiring, and recidivism predi
arXiv:2605.15491v1 Announce Type: cross Abstract: Layer pruning removes entire Transformer decoder blocks from large language models, but introduces a mismatch between the hidden state received by the
arXiv:2602.20207v3 Announce Type: replace-cross Abstract: Knowledge editing in Large Language Models (LLMs) aims to update the model's prediction for a specific query to a desired target while preserv
arXiv:2605.15290v1 Announce Type: cross Abstract: Hyperparameter transfer across model architectures dramatically reduces the amount of compute necessary for tuning large language models (LLMs). The m
arXiv:2605.15250v1 Announce Type: cross Abstract: Multi-head Latent Attention (MLA), the attention used in DeepSeek-V2/V3, jointly compresses keys and values into a low-rank latent and matches the H10
arXiv:2512.06655v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity obj