AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlog
85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Research

What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression

DGX agent

arXiv:2606.01292v1 Announce Type: cross Abstract: Teacher-Student Knowledge Transfer (KT) is ubiquitous in modern machine learning, ranging from classical model compression via Knowledge Distillation

researcharxiv-cs-ai
2 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

DGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

DGX agent

arXiv:2602.08236v2 Announce Type: replace-cross Abstract: Despite rapid progress in MLLMs, visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under uns

model-releasesarxiv-cs-ai
2 Jun 2026
Research

When Data Is Scarce: Scaling Sparse Language Models with Repeated Training

DGX agent

arXiv:2606.01155v1 Announce Type: cross Abstract: Scaling laws for dense LLMs under infinite data are well explored, but how sparsity interacts with limited data is not. In this work, we study sparse

researcharxiv-cs-ai
2 Jun 2026
Research

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

DGX agent

arXiv:2606.02378v1 Announce Type: cross Abstract: We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (de

researcharxiv-cs-ai
2 Jun 2026
Safety

When Does Predictive Inverse Dynamics Outperform Behavior Cloning?

DGX agent

arXiv:2601.21718v2 Announce Type: replace-cross Abstract: Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent work

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

When Jokes Cross the Line: Analyzing Regular Humor and Dark Humor in YouTube Shorts

DGX agent

arXiv:2606.00046v1 Announce Type: cross Abstract: Video platforms such as YouTube have reshaped how users engage with entertainment and information, emphasizing brief, highly engaging content such as

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

DGX agent

arXiv:2606.00448v1 Announce Type: cross Abstract: LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agen

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs

DGX agent

arXiv:2602.03554v2 Announce Type: replace-cross Abstract: Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evalu

model-releasesarxiv-cs-ai
2 Jun 2026
Research

When Softmax Fails at the Top: Extreme Value Corrections for InfoNCE

DGX agent

arXiv:2606.00262v1 Announce Type: cross Abstract: InfoNCE is the standard contrastive learning objective, but its softmax form is not only a computational convenience: it also encodes a statistical as

researcharxiv-cs-ai
2 Jun 2026
Model Releases

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

DGX agent

arXiv:2606.02060v1 Announce Type: new Abstract: Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final ans

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

DGX agent

arXiv:2606.02255v1 Announce Type: cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who p

researcharxiv-cs-ai
2 Jun 2026
Safety

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

DGX agent

arXiv:2511.05613v2 Announce Type: replace-cross Abstract: Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risk

safetyarxiv-cs-ai
2 Jun 2026
Applications

Why Do Time Series Models Need Long Context Windows?

DGX agent

arXiv:2606.01999v1 Announce Type: cross Abstract: Modern deep learning models for forecasting groups of time series rely on increasingly longer observation windows. However, the benefit of increasing

applicationsarxiv-cs-ai
2 Jun 2026
Model Releases

Why Not Hyperparameter-Friendly Optimisation? A Monotonic Adaptive Norm Rescaling Approach For Long-Tailed Recognition

DGX agent

arXiv:2606.02526v1 Announce Type: cross Abstract: Long-tailed recognition poses a significant challenge for deep learning. The two-stage decoupling paradigm, which separates representation learning fr

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

DGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

XAI-SOH-FL: Enhancing SOH-FL with Adaptive Aggregation and Explainable AI for Intrusion Detection in Heterogeneous IoT

DGX agent

arXiv:2606.00134v1 Announce Type: cross Abstract: Intrusion Detection Systems (IDS) in Internet of Things (IoT) environments face significant challenges due to data heterogeneity, lack of labeled data

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

You Can Learn Tokenization End-to-End with Reinforcement Learning

DGX agent

arXiv:2602.13940v2 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend t

model-releasesarxiv-cs-ai
2 Jun 2026
Tutorials

You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models

DGX agent

arXiv:2603.00133v2 Announce Type: replace-cross Abstract: Generative models have been shown to 'memorize' certain training data, leading to verbatim or near-verbatim generating images, which may cause

tutorialsarxiv-cs-ai
2 Jun 2026
Model Releases

Zamba2-VL Technical Report

DGX agent

arXiv:2606.00390v1 Announce Type: cross Abstract: We present Zamba2-VL, a suite of vision-language models built on Zamba2, a hybrid language-model architecture combining Mamba2 state-space layers with

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Zero-Shot Off-Policy Learning

DGX agent

arXiv:2602.01962v2 Announce Type: replace-cross Abstract: Off-policy learning methods seek to derive an optimal policy directly from a fixed dataset of prior interactions. This objective presents sign

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

DGX agent

arXiv:2602.08964v2 Announce Type: replace-cross Abstract: Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals

agentsarxiv-cs-ai
1 Jun 2026
Model Releases

A Kinetic Energy Perspective of Flow Matching

DGX agent

arXiv:2602.07928v2 Announce Type: replace-cross Abstract: Flow-based generative models can be viewed through a physics lens: sampling transports a particle from noise to data by integrating a learned

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

A Novel Global Context-aware Deep Neural Network for Enhanced Brain Tumor Segmentation using Magnetic Resonance Images

DGX agent

arXiv:2605.30510v1 Announce Type: cross Abstract: Brain cancer's severity necessitates precise brain tumor segmentation, which is crucial for effective brain tumor diagnosis. Manual identification, bu

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI

DGX agent

arXiv:2605.31021v1 Announce Type: new Abstract: Current alignment paradigms for generative artificial intelligence rely predominantly on monolithic benchmarking frameworks that reduce the plurality of

safetyarxiv-cs-ai
1 Jun 2026
Research

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models

DGX agent

arXiv:2605.31080v1 Announce Type: cross Abstract: Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy

researcharxiv-cs-ai
1 Jun 2026
Agents

A Unified and Reproducible Experimentation Framework for Speech Understanding

DGX agent

arXiv:2605.30899v1 Announce Type: cross Abstract: Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable eva

agentsarxiv-cs-ai
1 Jun 2026
Safety

A Unified Framework for Gradient Aggregation in Multi-Objective Optimization

DGX agent

arXiv:2605.30452v1 Announce Type: cross Abstract: Many machine learning problems involve multiple inherent trade-offs that are best addressed by gradient-based multi-objective optimization (MOO) algor

safetyarxiv-cs-ai
1 Jun 2026
Safety

Active Timepoint Selection for Learning Measure-Valued Trajectories

DGX agent

arXiv:2605.30625v1 Announce Type: cross Abstract: Inferring continuous probability paths from sparse snapshots is a fundamental challenge in domains like single-cell biology, where high-fidelity data

safetyarxiv-cs-ai
1 Jun 2026
Safety

AI Loss of Control Incident Management: Response & Resilience

DGX agent

arXiv:2605.30406v1 Announce Type: cross Abstract: Recent research demonstrating AI systems exhibiting deception and shutdown resistance suggests that AI loss of control (LOC) is an urgent policy conce

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

AMix-2: Establishing Protein as a Native Modality in Large Language Models

DGX agent

arXiv:2605.30963v1 Announce Type: cross Abstract: We present AMix-2, a protein-text foundation model that establishes protein as a native modality in large language models (LLMs), unifying protein und

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

An Odd Estimator for Shapley Values

DGX agent

arXiv:2602.01399v2 Announce Type: replace-cross Abstract: The Shapley value is a ubiquitous framework for attribution in machine learning, encompassing feature importance, data valuation, and causal i

model-releasesarxiv-cs-ai
1 Jun 2026
Local Ai

An Organization-Scoped LLM Agent Runtime Architecture for Regulated Cybersecurity Operations

DGX agent

arXiv:2605.30604v1 Announce Type: cross Abstract: Regulated cybersecurity workflows lack a runtime substrate that enforces organization-level scope across retrieval, tool calls, memory, findings, repo

local-aiarxiv-cs-ai
1 Jun 2026
Research

AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

DGX agent

arXiv:2605.31053v1 Announce Type: cross Abstract: Controllable music editing is to modify high-level attributes while strictly preserving rhythmic and melodic structures. However, this task is challen

researcharxiv-cs-ai
1 Jun 2026
Safety

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

DGX agent

arXiv:2605.31034v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling

safetyarxiv-cs-ai
1 Jun 2026
Agents

Answer-Set-Programming-based Abstractions for Reinforcement Learning

DGX agent

arXiv:2605.31444v1 Announce Type: new Abstract: Reinforcement Learning (RL) enables autonomous agents to learn policies from experience, but realistic problems often involve enormous state spaces, mak

agentsarxiv-cs-ai
1 Jun 2026
Tutorials

Appropriateness of Empathy in AI: A Signal-Cost Perspective

DGX agent

arXiv:2605.31340v1 Announce Type: cross Abstract: The appropriateness of empathy in AI has emerged as a critical concern, as excessive empathy risks seeming manipulative while insufficient empathy app

tutorialsarxiv-cs-ai
1 Jun 2026
Model Releases

Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery

DGX agent

arXiv:2502.15224v2 Announce Type: replace-cross Abstract: Interactive discovery requires agents to maintain and update structured beliefs over many rounds of feedback. Before evaluating agents in nois

model-releasesarxiv-cs-ai
1 Jun 2026
Agents

Automatically Attacking Software Reverse Engineering AI Agents

DGX agent

arXiv:2605.30667v1 Announce Type: cross Abstract: Software tools for reverse engineering executable binary files, such as Ghidra, enable malware analysts to safely conduct robust static analysis witho

agentsarxiv-cs-ai
1 Jun 2026
Research

Autoregressive Visual Generation Needs a Prologue

DGX agent

arXiv:2605.06137v2 Announce Type: replace-cross Abstract: In this work, we propose Prologue, an approach to bridging the reconstruction-generation gap in autoregressive (AR) image generation. Instead

researcharxiv-cs-ai
1 Jun 2026
Agents

AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle

DGX agent

arXiv:2605.31468v1 Announce Type: new Abstract: Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review

agentsarxiv-cs-ai
1 Jun 2026
Model Releases

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

DGX agent

arXiv:2605.31212v1 Announce Type: cross Abstract: AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully rep

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

Benchmarking Machine Learning Uncertainty Quantification Methodologies for Predicting Turbine Gas Temperature Degradation

DGX agent

arXiv:2605.30585v1 Announce Type: cross Abstract: Effective prognostics and health management of modern engines relies on accurate turbine gas temperature predictions and robust uncertainty quantifica

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage

DGX agent

arXiv:2605.30826v1 Announce Type: cross Abstract: Biomedical NER is deceptively simple for modern LLMs: plausible biomedical mentions are easy to surface, but corpus-convention correctness depends on

model-releasesarxiv-cs-ai
1 Jun 2026
Research

Beyond Classification: Dynamic Adapter Routing for Continual Multimodal Retrieval

DGX agent

arXiv:2605.31229v1 Announce Type: cross Abstract: While retrieval is a core function of vision-language models, continually updating these models for retrieval tasks remains critically underexplored.

researcharxiv-cs-ai
1 Jun 2026
Applications

Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions

DGX agent

arXiv:2501.04661v3 Announce Type: replace-cross Abstract: The web-scale of pretraining data has created an important evaluation challenge: to disentangle linguistic competence on cases well-represente

applicationsarxiv-cs-ai
1 Jun 2026
Safety

Biases in the Blind Spot: Detecting What LLMs Fail to Mention

DGX agent

arXiv:2602.10117v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We cal

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs

DGX agent

arXiv:2605.30900v1 Announce Type: new Abstract: Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move an

model-releasesarxiv-cs-ai
1 Jun 2026
← Previous
1…232233234235236…452
Next →