AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
Research

LPT: Less-overfitting Prompt Tuning for Vision-Language Model

DGX agent

arXiv:2410.10247v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated exceptional generalization capabilities for downstream tasks. Due to its efficiency, prompt le

researcharxiv-cs-ai
12 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models

DGX agent

arXiv:2605.09879v1 Announce Type: new Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misal

model-releasesarxiv-cs-ai
12 May 2026
Safety

M^3: Reframing Training Measures for Discretized Physical Simulations

DGX agent

arXiv:2605.08843v1 Announce Type: new Abstract: Neural surrogate models for physical simulations are trained on discretized samples of continuous domains, where the induced empirical measure leads to

safetyarxiv-cs-ai
12 May 2026
Model Releases

MaD Physics: Evaluating information seeking under constraints in physical environments

DGX agent

arXiv:2605.10820v1 Announce Type: new Abstract: Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of mea

model-releasesarxiv-cs-ai
12 May 2026
Agents

MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs

DGX agent

arXiv:2605.10064v1 Announce Type: new Abstract: Self-evolving language-model agents must decide what to learn next and how to preserve what they have learned across iterations. Existing systems typica

agentsarxiv-cs-ai
12 May 2026
Model Releases

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

DGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

model-releasesarxiv-cs-ai
12 May 2026
Safety

MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service

DGX agent

arXiv:2605.08527v1 Announce Type: cross Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has significantly improved the reasoning capabilities of large language models (LLMs), particula

safetyarxiv-cs-ai
12 May 2026
Applications

Marrying Generative Model of Healthcare Events with Digital Twin of Social Determinants of Health for Disease Reasoning

DGX agent

arXiv:2605.09771v1 Announce Type: new Abstract: Despite the central role of sensor-derived measurements such as imaging traits and plasma biomarkers in biomedical research and clinical practice, exist

applicationsarxiv-cs-ai
12 May 2026
Research

Matching Meaning at Scale: Evaluating Semantic Search for 18th-Century Intellectual History through the Case of Locke

DGX agent

arXiv:2605.09236v1 Announce Type: cross Abstract: While digitized corpora have transformed the study of intellectual transmission, current methods rely heavily on lexical text reuse detection, capturi

researcharxiv-cs-ai
12 May 2026
Model Releases

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

DGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Agents

MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study

DGX agent

arXiv:2605.10763v1 Announce Type: new Abstract: LLMs are increasingly deployed as autonomous agents with access to tools, databases, and external services, yet practitioners (across different sectors)

agentsarxiv-cs-ai
12 May 2026
Local Ai

Mazocarta: A Seeded Procedural Deckbuilder for Instrumented Game Development

DGX agent

arXiv:2605.08319v1 Announce Type: cross Abstract: Mazocarta is a seeded procedural tactical deckbuilder implemented in Rust, compiled to WebAssembly for browser play, and executable natively for simul

local-aiarxiv-cs-ai
12 May 2026
Model Releases

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

DGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

DGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC^2: Monte Carlo Correction for Fast Elliptic PDE Solving

DGX agent

arXiv:2605.09288v1 Announce Type: cross Abstract: Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solv

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

DGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDGYM: Benchmarking AI Agents on Molecular Simulations

DGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

DGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

model-releasesarxiv-cs-ai
12 May 2026
Research

Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings

DGX agent

arXiv:2605.10606v1 Announce Type: cross Abstract: Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embedd

researcharxiv-cs-ai
12 May 2026
Model Releases

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

DGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

DGX agent

arXiv:2507.23511v3 Announce Type: replace-cross Abstract: While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. Th

model-releasesarxiv-cs-ai
12 May 2026
Safety

Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

DGX agent

arXiv:2605.08426v1 Announce Type: cross Abstract: Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI s

safetyarxiv-cs-ai
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Applications

Medical Model Synthesis Architectures: A Case Study

DGX agent

arXiv:2605.09716v1 Announce Type: new Abstract: Medicine is rife with high-stakes uncertainty. Doctors routinely make clinical judgments and decisions that juggle many fundamental unknowns, like predi

applicationsarxiv-cs-ai
12 May 2026
Model Releases

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

DGX agent

arXiv:2605.09661v1 Announce Type: cross Abstract: Large language models (LLMs) have saturated standard medical benchmarks that test factual recall, yet their ability to perform higher-order reasoning,

model-releasesarxiv-cs-ai
12 May 2026
Research

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction

DGX agent

arXiv:2605.08094v1 Announce Type: cross Abstract: Accurate clinical diagnosis requires extensive domain knowledge and complex clinical reasoning capabilities. Although large language models (LLMs) hol

researcharxiv-cs-ai
12 May 2026
Research

Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

DGX agent

arXiv:2605.09270v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization.

researcharxiv-cs-ai
12 May 2026
Safety

Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

DGX agent

arXiv:2605.06225v2 Announce Type: replace-cross Abstract: Steering large language models (LLMs) is usually done by either instruction prompting or activation steering. Prompting often gives strong con

safetyarxiv-cs-ai
12 May 2026
Model Releases

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

DGX agent

arXiv:2605.08374v1 Announce Type: new Abstract: Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval

model-releasesarxiv-cs-ai
12 May 2026
Agents

MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading

DGX agent

arXiv:2605.10268v1 Announce Type: cross Abstract: To tackle long-context reasoning tasks without the quadratic complexity of standard attention mechanisms, approaches based on agent memory have emerge

agentsarxiv-cs-ai
12 May 2026
Safety

Mental Health AI Safety Claims Must Preserve Temporal Evidence

DGX agent

arXiv:2605.08827v1 Announce Type: new Abstract: The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, o

safetyarxiv-cs-ai
12 May 2026
Safety

MePo: Meta Post-Refinement for Rehearsal-Free General Continual Learning

DGX agent

arXiv:2602.07940v3 Announce Type: replace Abstract: To cope with uncertain changes of the external world, intelligent systems must continually learn from complex, evolving environments and respond in

safetyarxiv-cs-ai
12 May 2026
Model Releases

MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups

DGX agent

arXiv:2603.13452v2 Announce Type: replace Abstract: Fairness in machine learning is predominantly evaluated through outcome-oriented metrics, such as Demographic parity, which measure whether predicti

model-releasesarxiv-cs-ai
12 May 2026
Research

MeshFIM: Local Low-Poly Mesh Editing via Fill-in-the-Middle Autoregressive Generation

DGX agent

arXiv:2605.08744v1 Announce Type: cross Abstract: Autoregressive (AR) models can generate high-quality low-poly meshes from point clouds, but they still operate in an all-or-nothing manner: when a loc

researcharxiv-cs-ai
12 May 2026
Research

Metacognitive Behavioral Tuning of Large Language Models for Multi-Hop Question Answering

DGX agent

arXiv:2602.22508v2 Announce Type: replace Abstract: Large Language Models (LLMs) often produce incorrect answers on multi-hop question answering even when the reasoning trace already contains a correc

researcharxiv-cs-ai
12 May 2026
Model Releases

Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

DGX agent

arXiv:2605.09708v1 Announce Type: cross Abstract: We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

DGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

model-releasesarxiv-cs-ai
12 May 2026
Hardware

mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters

DGX agent

arXiv:2605.08300v1 Announce Type: cross Abstract: Manifold-Constrained Hyper-Connections (mHC) introduce a stability-motivated variant of multi stream residual mixing by constraining residual stream m

hardwarearxiv-cs-ai
12 May 2026
Research

Micro-Defects Expose Macro-Fakes: Detecting AI-Generated Images via Local Distributional Shifts

DGX agent

arXiv:2605.09296v1 Announce Type: cross Abstract: Recent generative models can produce images that appear highly realistic, raising challenges in distinguishing real and AI-generated images. Yet exist

researcharxiv-cs-ai
12 May 2026
Model Releases

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

DGX agent

arXiv:2605.10120v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as micr

model-releasesarxiv-cs-ai
12 May 2026
Safety

Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models

DGX agent

arXiv:2605.08472v1 Announce Type: new Abstract: The effectiveness of Reinforcement Learning (RL) in Large Language Models (LLMs) depends on the nature and diversity of the data used before and during

safetyarxiv-cs-ai
12 May 2026
Model Releases

MIDUS: Memory-Infused Depth Up-Scaling

DGX agent

arXiv:2512.13751v2 Announce Type: replace-cross Abstract: Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scalin

model-releasesarxiv-cs-ai
12 May 2026
Agents

MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction

DGX agent

arXiv:2605.08670v1 Announce Type: new Abstract: Large language model (LLM) powered AI agents have emerged as a promising paradigm for autonomous problem-solving, yet they continue to struggle with com

agentsarxiv-cs-ai
12 May 2026
Research

MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

DGX agent

arXiv:2603.13131v3 Announce Type: replace Abstract: Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A centra

researcharxiv-cs-ai
12 May 2026
Model Releases

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

DGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

model-releasesarxiv-cs-ai
12 May 2026
Safety

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

DGX agent

arXiv:2605.08277v1 Announce Type: cross Abstract: Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demon

safetyarxiv-cs-ai
12 May 2026
Research

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection

DGX agent

arXiv:2507.07871v4 Announce Type: replace-cross Abstract: Watermarking enables GenAI providers to verify whether content was generated by their models. A watermark is a hidden signal in the content, w

researcharxiv-cs-ai
12 May 2026
← Previous
1…337338339340341…448
Next →