AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
20,981 results
11 Aug 2026

MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models

Model ReleasesDGX agent

arXiv:2608.08503v1 Announce Type: new Abstract: Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) su

Matryoshka Language Model Suites

Model ReleasesDGX agent

arXiv:2608.09703v1 Announce Type: new Abstract: Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inferen

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks

Model ReleasesDGX agent

arXiv:2507.19634v4 Announce Type: replace-cross Abstract: Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a s


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Model ReleasesDGX agent

arXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.

Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets

ResearchDGX agent

arXiv:2512.01045v2 Announce Type: replace Abstract: Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets, yet the

MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning

SafetyDGX agent

arXiv:2608.08623v1 Announce Type: new Abstract: In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated usi

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

Model ReleasesDGX agent

arXiv:2608.09818v1 Announce Type: cross Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-

MELLON - Multimodal Enhanced LLM for Online Navigation

Model ReleasesDGX agent

arXiv:2608.09121v1 Announce Type: new Abstract: Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or

Mendel Godel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution

AgentsDGX agent

arXiv:2608.07645v1 Announce Type: new Abstract: Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing

Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality

HardwareDGX agent

arXiv:2512.20968v2 Announce Type: replace-cross Abstract: Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited paralle

Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses

ApplicationsDGX agent

arXiv:2608.07946v1 Announce Type: cross Abstract: Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic identifiers,

Metanormative Theory for RL-Based Moral Agents

SafetyDGX agent

arXiv:2608.08220v1 Announce Type: new Abstract: The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and

MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents

Model ReleasesDGX agent

arXiv:2608.07533v1 Announce Type: new Abstract: An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodied agents pri

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

SafetyDGX agent

arXiv:2608.09836v1 Announce Type: new Abstract: On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement,

Mitigating Gender Bias in English to Romanian Machine Translation

SafetyDGX agent

arXiv:2608.08606v1 Announce Type: cross Abstract: Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a

Mitigating Over-Personalization in LLMs via Structured Memory

ResearchDGX agent

arXiv:2608.08300v1 Announce Type: new Abstract: Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions. However, when stored user informati

MixFormer: Linear Transformer with Mixture of Memory Experts

ResearchDGX agent

arXiv:2608.09468v1 Announce Type: cross Abstract: State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in l

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

Model ReleasesDGX agent

arXiv:2608.09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information e

Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models

Model ReleasesDGX agent

arXiv:2608.09696v1 Announce Type: new Abstract: Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a mechanistic, causal model, not a c

Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization

Local AiDGX agent

arXiv:2608.09801v1 Announce Type: cross Abstract: Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a m

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

Model ReleasesDGX agent

arXiv:2603.28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mis

MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries

ResearchDGX agent

arXiv:2608.09764v1 Announce Type: cross Abstract: Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatial observati

Monotonicity-Guided Bottom-Up Petri Net Discovery: The SPECpp Framework

ResearchDGX agent

arXiv:2608.09398v1 Announce Type: cross Abstract: Process discovery is one of the central challenges in process mining. Petri nets are particularly attractive because simple local constructs can expre

MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts

Model ReleasesDGX agent

arXiv:2608.09251v1 Announce Type: cross Abstract: Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly

MOSAIC: Adversarial Co-evolution of Specialist Heuristics and Problem Instances for LLM-based Automated Heuristic Design

Model ReleasesDGX agent

arXiv:2608.07544v1 Announce Type: cross Abstract: Automated heuristic design (AHD) with large language models (LLMs) has produced strong heuristics for combinatorial optimization problems (COPs). Yet

Motif 3: Technical Report

SafetyDGX agent

arXiv:2608.09119v1 Announce Type: new Abstract: We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each spar

Multi-Agent AI Safety as an Institutional Design Problem

SafetyDGX agent

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent wo

Multi-agent discovery of practical quantum LDPC codes

AgentsDGX agent

arXiv:2608.08996v1 Announce Type: cross Abstract: Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful finite-length i

Multi-Branch Policy Optimization for Multimodal Large Language Models

SafetyDGX agent

arXiv:2608.07581v1 Announce Type: cross Abstract: Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a si

Multi-Granular Node Pruning for Causal Circuit Discovery

TutorialsDGX agent

arXiv:2512.10903v3 Announce Type: replace Abstract: Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existing approac

Multi-Task Consistency-based Detection of Adversarial Attacks

AgentsDGX agent

arXiv:2608.07750v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to adversarial attac

Multilingual Agent-Based World Modeling for Social Science

Model ReleasesDGX agent

arXiv:2512.07195v2 Announce Type: replace-cross Abstract: Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly mon

Multimodal Federated Learning under Dual-Axis Modality Missingness

Local AiDGX agent

arXiv:2608.09240v1 Announce Type: cross Abstract: Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic deployments

Multimodal Model Diffing for Feature Discovery and Control

SafetyDGX agent

arXiv:2608.09928v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

ResearchDGX agent

arXiv:2608.09035v1 Announce Type: cross Abstract: Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of

NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs

ResearchDGX agent

arXiv:2608.08107v1 Announce Type: cross Abstract: Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired duri

NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2510.18940v2 Announce Type: replace-cross Abstract: Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. T

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

SafetyDGX agent

arXiv:2509.03985v2 Announce Type: replace-cross Abstract: In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Howev

Neuronal Attention Circuit (NAC) for Representation Learning

AgentsDGX agent

arXiv:2512.10282v4 Announce Type: replace Abstract: Attention improves representation learning over RNNs, but its discrete nature limits continuous-time (CT) modeling. We introduce Neuronal Attention

NeuroPilot: An Agent-Driven Smart Pipeline for Processing, Quality Control, and Managing Neuroimages

AgentsDGX agent

arXiv:2608.07541v1 Announce Type: cross Abstract: Transforming raw neuroimage archives into analysis-ready derivatives relies on three brittle stages: data standardization, modality-specific preproces

NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation

AgentsDGX agent

arXiv:2608.09636v1 Announce Type: cross Abstract: Accurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of neurons poses

Neurosymbolic Discovery of Algebraic Graph Constructions

Model ReleasesDGX agent

arXiv:2608.08118v1 Announce Type: new Abstract: There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. These methods return the

NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

Model ReleasesDGX agent

arXiv:2608.07530v1 Announce Type: new Abstract: SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that m

Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression

Model ReleasesDGX agent

arXiv:2608.09176v1 Announce Type: cross Abstract: Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize

Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection

AgentsDGX agent

arXiv:2608.08939v1 Announce Type: new Abstract: The rise of autonomous AI agents represents a major paradigm shift in how users interact with mobile devices. Frameworks such as MobileRun and Mobile-Us

Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents

AgentsDGX agent

arXiv:2608.08389v1 Announce Type: new Abstract: Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the margina

OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents

Model ReleasesDGX agent

arXiv:2608.08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new

Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models

Model ReleasesDGX agent

arXiv:2608.09227v1 Announce Type: new Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally pr

On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation

Local AiDGX agent

arXiv:2608.08566v1 Announce Type: cross Abstract: Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce. Automated diagnosis based on microsc

On the Robustness of LLMs' Internal Representation of Code Correctness

ResearchDGX agent

arXiv:2608.08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as r

One Adapter Pair per Model: A Universal Activation Interface for Language Models

Model ReleasesDGX agent

arXiv:2608.09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to

Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models

Model ReleasesDGX agent

arXiv:2608.09666v1 Announce Type: new Abstract: Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands sampling hun

Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects

Model ReleasesDGX agent

arXiv:2608.07577v1 Announce Type: cross Abstract: A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-drawn carri

Open-World Semantic Segmentation with Sensitivity Modeling

ResearchDGX agent

arXiv:2608.08308v1 Announce Type: cross Abstract: Modern vision systems must operate in 'open-world' settings, where models must recognize known categories and detect unseen or anomalous content. Conv

OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks

Model ReleasesDGX agent

arXiv:2608.09380v1 Announce Type: new Abstract: Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools, verify results, and recover from failures in cont

OpenMHC: Accelerating the Science of Wearable Foundation Models

Model ReleasesDGX agent

arXiv:2607.16235v3 Announce Type: replace-cross Abstract: Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. However,

Optimal Transport for Machine Learners

ResearchDGX agent

arXiv:2505.06589v3 Announce Type: replace-cross Abstract: Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-condit

Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution

Model ReleasesDGX agent

arXiv:2608.08311v1 Announce Type: cross Abstract: We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits

P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization

Local AiDGX agent

arXiv:2608.07549v1 Announce Type: cross Abstract: Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampl

P^{3}: Joint Program-and-Proof Planning for Verified Code Generation

Model ReleasesDGX agent

arXiv:2608.09277v1 Announce Type: new Abstract: Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the program meets a

← Previous
1…910111213…350
Next →