AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
12 May 2026

LENS: LLM-Enabled Narrative Synthesis for Mental Health by Aligning Multimodal Sensing with Language Models

ResearchDGX agent

arXiv:2512.23025v2 Announce Type: replace-cross Abstract: Multimodal health sensing offers rich behavioral signals for assessing mental health, yet translating these numerical time-series measurements

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

Model ReleasesDGX agent

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

Leveraging LLMs to Automate Energy-Aware Refactoring of Parallel Scientific Codes

HardwareDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2505.02184v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for generating parallel scientific codes, with a primary focus on generating functionally correct

LEVI: Stronger Search Architectures Can Substitute for Larger LLMs in Evolutionary Search

Model ReleasesDGX agent

arXiv:2605.09764v1 Announce Type: cross Abstract: LLM-guided evolutionary methods such as AlphaEvolve have proven effective in domains like math, systems research, and algorithmic discovery, but their

LILO: Bayesian Optimization with Natural Language Feedback

ApplicationsDGX agent

arXiv:2510.17671v2 Announce Type: replace-cross Abstract: Many real-world optimization problems are guided by complex, subjective preferences that are difficult to express as explicit closed-form obje

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation

ApplicationsDGX agent

arXiv:2605.10593v1 Announce Type: new Abstract: We demonstrate LLARS (LLM Assisted Research System), an open-source platform that bridges the gap between domain experts and developers for building LLM

LLaVA-CKD: Bottom-Up Cascaded Knowledge Distillation for Vision-Language Models

ApplicationsDGX agent

arXiv:2605.10641v1 Announce Type: cross Abstract: Large Vision-Language Models (VLMs) are successful in addressing a multitude of vision-language understanding tasks, such as Visual Question Answering

LLM Advertisement based on Neuron Auctions

SafetyDGX agent

arXiv:2605.08326v1 Announce Type: cross Abstract: As Large Language Models (LLMs) transition into conversational agents, generative advertising emerges as a crucial monetization strategy. However, emb

LLM-Agnostic Semantic Representation Attack

SafetyDGX agent

arXiv:2605.08898v1 Announce Type: cross Abstract: Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent t

LLM-Augmented Chemical Synthesis and Design Decision Programs

TutorialsDGX agent

arXiv:2505.07027v2 Announce Type: replace Abstract: Retrosynthesis, the process of breaking down a target molecule into simpler precursors through a series of valid reactions, stands at the core of or

LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers

ResearchDGX agent

arXiv:2503.14434v3 Announce Type: replace-cross Abstract: Automated feature engineering plays a critical role in improving predictive model performance for tabular learning tasks. Traditional automate

LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs

Local AiDGX agent

arXiv:2605.09542v1 Announce Type: new Abstract: Extracting multi-step explanations from knowledge graphs poses a combinatorial challenge requiring both heuristic guidance (as candidates proliferate wi

LLM-guided Semi-Supervised Approaches for Social Media Crisis Data Classification

ApplicationsDGX agent

arXiv:2605.08448v1 Announce Type: new Abstract: Semi-supervised learning approaches have been investigated as a means to enhance the analysis of social media data in disaster management contexts. In t

LLM Jaggedness Unlocks Scientific Creativity

Model ReleasesDGX agent

arXiv:2605.10574v1 Announce Type: new Abstract: As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing uneven

LLM Translation of Compiler Intermediate Representation

Model ReleasesDGX agent

arXiv:2605.08247v1 Announce Type: cross Abstract: GCC and LLVM underpin much of modern software infrastructure, relying on distinct Intermediate Representations (IRs) to drive optimizations and code g

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight

Model ReleasesDGX agent

arXiv:2605.08321v1 Announce Type: cross Abstract: LLMs are increasingly capable of persuasion, which raises the question of how to protect users against manipulation. In a preregistered user study (N=

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

Model ReleasesDGX agent

arXiv:2605.10401v1 Announce Type: new Abstract: Efficient branching policies are essential for accelerating Mixed Integer Linear Programming (MILP) solvers. Their design has long relied on hand-crafte

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

Model ReleasesDGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

Log analysis is necessary for credible evaluation of AI agents

Model ReleasesDGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

HardwareDGX agent

arXiv:2605.10886v1 Announce Type: cross Abstract: Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language

Long-Horizon Q-Learning: Accurate Value Learning via n-Step Inequalities

SafetyDGX agent

arXiv:2605.05812v2 Announce Type: replace Abstract: Off-policy, value-based reinforcement learning methods such as Q-learning are appealing because they can learn from arbitrary experience, including

LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models

SafetyDGX agent

arXiv:2605.09948v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal for action p

Lost in Translation? Exploring the Shift in Grammatical Gender from Latin to Occitan

ResearchDGX agent

arXiv:2605.09156v1 Announce Type: cross Abstract: The diachronic evolution from Latin to the Romance languages involved a restructuring of the grammatical gender system from a tripartite configuration

LPT: Less-overfitting Prompt Tuning for Vision-Language Model

ResearchDGX agent

arXiv:2410.10247v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated exceptional generalization capabilities for downstream tasks. Due to its efficiency, prompt le

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09879v1 Announce Type: new Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misal

M^3: Reframing Training Measures for Discretized Physical Simulations

SafetyDGX agent

arXiv:2605.08843v1 Announce Type: new Abstract: Neural surrogate models for physical simulations are trained on discretized samples of continuous domains, where the induced empirical measure leads to

MaD Physics: Evaluating information seeking under constraints in physical environments

Model ReleasesDGX agent

arXiv:2605.10820v1 Announce Type: new Abstract: Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of mea

MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs

AgentsDGX agent

arXiv:2605.10064v1 Announce Type: new Abstract: Self-evolving language-model agents must decide what to learn next and how to preserve what they have learned across iterations. Existing systems typica

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Model ReleasesDGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service

SafetyDGX agent

arXiv:2605.08527v1 Announce Type: cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has significantly improved the reasoning capabilities of large language models (LLMs), particula

Marrying Generative Model of Healthcare Events with Digital Twin of Social Determinants of Health for Disease Reasoning

ApplicationsDGX agent

arXiv:2605.09771v1 Announce Type: new Abstract: Despite the central role of sensor-derived measurements such as imaging traits and plasma biomarkers in biomedical research and clinical practice, exist

Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke

ResearchDGX agent

arXiv:2605.09236v1 Announce Type: cross Abstract: While digitized corpora have transformed the study of intellectual transmission, current methods rely heavily on lexical text reuse detection, capturi

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

Model ReleasesDGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

Model ReleasesDGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study

AgentsDGX agent

arXiv:2605.10763v1 Announce Type: new Abstract: LLMs are increasingly deployed as autonomous agents with access to tools, databases, and external services, yet practitioners (across different sectors)

Mazocarta: A Seeded Procedural Deckbuilder for Instrumented Game Development

Local AiDGX agent

arXiv:2605.08319v1 Announce Type: cross Abstract: Mazocarta is a seeded procedural tactical deckbuilder implemented in Rust, compiled to WebAssembly for browser play, and executable natively for simul

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

Model ReleasesDGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

Model ReleasesDGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

MC^2: Monte Carlo Correction for Fast Elliptic PDE Solving

Model ReleasesDGX agent

arXiv:2605.09288v1 Announce Type: cross Abstract: Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solv

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

Model ReleasesDGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

MDGYM: Benchmarking AI Agents on Molecular Simulations

Model ReleasesDGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Model ReleasesDGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings

ResearchDGX agent

arXiv:2605.10606v1 Announce Type: cross Abstract: Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embedd

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

Model ReleasesDGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

Model ReleasesDGX agent

arXiv:2507.23511v3 Announce Type: replace-cross Abstract: While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. Th

Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

SafetyDGX agent

arXiv:2605.08426v1 Announce Type: cross Abstract: Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI s

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

Model ReleasesDGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

Medical Model Synthesis Architectures: A Case Study

ApplicationsDGX agent

arXiv:2605.09716v1 Announce Type: new Abstract: Medicine is rife with high-stakes uncertainty. Doctors routinely make clinical judgments and decisions that juggle many fundamental unknowns, like predi

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

Model ReleasesDGX agent

arXiv:2605.09661v1 Announce Type: cross Abstract: Large language models (LLMs) have saturated standard medical benchmarks that test factual recall, yet their ability to perform higher-order reasoning,

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction

ResearchDGX agent

arXiv:2605.08094v1 Announce Type: cross Abstract: Accurate clinical diagnosis requires extensive domain knowledge and complex clinical reasoning capabilities. Although large language models (LLMs) hol

Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

ResearchDGX agent

arXiv:2605.09270v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization.

Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

SafetyDGX agent

arXiv:2605.06225v2 Announce Type: replace-cross Abstract: Steering large language models (LLMs) is usually done by either instruction prompting or activation steering. Prompting often gives strong con

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

Model ReleasesDGX agent

arXiv:2605.08374v1 Announce Type: new Abstract: Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval

MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading

AgentsDGX agent

arXiv:2605.10268v1 Announce Type: cross Abstract: To tackle long-context reasoning tasks without the quadratic complexity of standard attention mechanisms, approaches based on agent memory have emerge

Mental Health AI Safety Claims Must Preserve Temporal Evidence

SafetyDGX agent

arXiv:2605.08827v1 Announce Type: new Abstract: The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, o

MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning

SafetyDGX agent

arXiv:2602.07940v3 Announce Type: replace Abstract: To cope with uncertain changes of the external world, intelligent systems must continually learn from complex, evolving environments and respond in

MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups

Model ReleasesDGX agent

arXiv:2603.13452v2 Announce Type: replace Abstract: Fairness in machine learning is predominantly evaluated through outcome-oriented metrics, such as Demographic parity, which measure whether predicti

MeshFIM: Local Low-Poly Mesh Editing via Fill-in-the-Middle Autoregressive Generation

ResearchDGX agent

arXiv:2605.08744v1 Announce Type: cross Abstract: Autoregressive (AR) models can generate high-quality low-poly meshes from point clouds, but they still operate in an all-or-nothing manner: when a loc

Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering

ResearchDGX agent

arXiv:2602.22508v2 Announce Type: replace Abstract: Large Language Models (LLMs) often produce incorrect answers on multi-hop question answering even when the reasoning trace already contains a correc

← Previous
1…269270271272273…358
Next →