AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Research

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

DGX agent

arXiv:2605.22012v1 Announce Type: new Abstract: Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasonin

researcharxiv-cs-cl
22 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?

DGX agent

arXiv:2510.07962v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated remarkable progress in reasoning, often through supervised fine-tuning (SFT). However, SFT is resourc

researcharxiv-cs-cl
22 May 2026
Model Releases

Linear Dynamics in the RLVR Training of Large Language Models

DGX agent

arXiv:2601.04537v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven significant performance gains in reasoning-oriented large language models (LL

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications

DGX agent

arXiv:2603.27355v2 Announce Type: replace-cross Abstract: We present a readiness harness for LLM and RAG applications that turns evaluation into a deployment decision workflow. The system combines aut

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

DGX agent

arXiv:2605.22177v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing f

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models

DGX agent

arXiv:2510.23090v2 Announce Type: replace Abstract: Recent advances have investigated the use of pretrained large language models (LLMs) for time-series forecasting by aligning numerical inputs with L

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

DGX agent

arXiv:2605.21796v1 Announce Type: cross Abstract: Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision

model-releasesarxiv-cs-cl
22 May 2026
Safety

Modeling Pathology-Like Behavioral Patterns in Language Models Through Behavioral Fine-Tuning

DGX agent

arXiv:2605.22356v1 Announce Type: new Abstract: Large language models are increasingly used as computational tools for modeling human-like behavior. We introduce a behavioral induction framework that

safetyarxiv-cs-cl
22 May 2026
Safety

Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora

DGX agent

arXiv:2605.22660v1 Announce Type: new Abstract: Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultur

safetyarxiv-cs-cl
22 May 2026
Research

More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts

DGX agent

arXiv:2605.22641v1 Announce Type: new Abstract: Detecting Schwartz values in political text is difficult because implicit cues often depend on surrounding arguments and fine-grained distinctions betwe

researcharxiv-cs-cl
22 May 2026
Model Releases

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

DGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

model-releasesarxiv-cs-cl
22 May 2026
Research

Multi-Stage Training for Abusive Comment Detection in Indic Languages

DGX agent

arXiv:2605.22380v1 Announce Type: new Abstract: In recent years social media has become an increasingly popular tool for communication. People use it to share their ideas, exchange information, and di

researcharxiv-cs-cl
22 May 2026
Safety

NaviAgent: Graph-Driven Bilevel Planning for Scalable Tool Orchestration

DGX agent

arXiv:2506.19500v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly act as function-call agents that invoke external tools to tackle tasks beyond their static knowledge

safetyarxiv-cs-cl
22 May 2026
Research

One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation

DGX agent

arXiv:2605.22544v1 Announce Type: new Abstract: Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point ev

researcharxiv-cs-cl
22 May 2026
Safety

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI

DGX agent

arXiv:2507.05660v3 Announce Type: replace-cross Abstract: Customizing Large Language Models (LLMs) on untrusted datasets poses severe risks of injecting toxic behaviors. In this work, we introduce Opt

safetyarxiv-cs-cl
22 May 2026
Research

Pattern-and-root inflectional morphology: the Arabic broken plural

DGX agent

arXiv:2605.22310v1 Announce Type: new Abstract: We present a substantially implemented model of description of the inflectional morphology of Arabic nouns, with special attention to the management of

researcharxiv-cs-cl
22 May 2026
Research

Planning in the LLM Era: Building for Reliability and Efficiency

DGX agent

arXiv:2605.21902v1 Announce Type: cross Abstract: Growing attention to intelligent agents has put a spotlight on one of their central capabilities: planning. Early attempts to leverage large language

researcharxiv-cs-cl
22 May 2026
Model Releases

Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation

DGX agent

arXiv:2605.22487v1 Announce Type: new Abstract: Recent advances in Multilingual Large Language Models (MLLMs) have significantly enhanced cross-lingual conversational capabilities, yet modeling cultur

model-releasesarxiv-cs-cl
22 May 2026
Research

Probabilistic Attribution For Large Language Models

DGX agent

arXiv:2605.21726v1 Announce Type: new Abstract: The generative nature of Large Language Models (LLMs) is reflected in the conditional probabilities they compute to sample each response token given the

researcharxiv-cs-cl
22 May 2026
Model Releases

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts

DGX agent

arXiv:2605.21776v1 Announce Type: new Abstract: Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether larg

model-releasesarxiv-cs-cl
22 May 2026
Agents

Psy-Chronicle:A Structured Pipeline for Synthesizing Long-Horizon Campus Psychological Counseling Dialogues

DGX agent

arXiv:2605.22140v1 Announce Type: new Abstract: In recent years, large language models have shown substantial potential in psychological support tasks. However, existing psychological counseling data

agentsarxiv-cs-cl
22 May 2026
Model Releases

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

DGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

model-releasesarxiv-cs-cl
22 May 2026
Applications

Quantizing Whisper-small: How design choices affect ASR performance

DGX agent

arXiv:2511.08093v2 Announce Type: replace-cross Abstract: Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computa

applicationsarxiv-cs-cl
22 May 2026
Model Releases

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

DGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

DGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

DGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

model-releasesarxiv-cs-cl
22 May 2026
Safety

Reducing Political Manipulation with Consistency Training

DGX agent

arXiv:2605.22771v1 Announce Type: new Abstract: Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from

safetyarxiv-cs-cl
22 May 2026
Agents

Reflecti-Mate: A Conversational Agent for Adaptive Decision-Making Support Through System 1 and System 2 Thinking

DGX agent

arXiv:2605.22509v1 Announce Type: cross Abstract: Making high-stakes personal decisions involves cognitive, emotional, and intuitive processes, and individuals differ in how they allocate attention ac

agentsarxiv-cs-cl
22 May 2026
Model Releases

Reflective Prompt Tuning through Language Model Function-Calling

DGX agent

arXiv:2605.21781v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Residual Skill Optimization for Text-to-SQL Ensembles

DGX agent

arXiv:2605.21792v1 Announce Type: new Abstract: Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded

model-releasesarxiv-cs-cl
22 May 2026
Research

Scene Abstraction for Lexical Semantics: Structured Representations of Situated Meaning

DGX agent

arXiv:2605.22542v1 Announce Type: new Abstract: Coffee and tea share many properties, yet they evoke strikingly different situations, atmospheres, and affective associations. These situated dimensions

researcharxiv-cs-cl
22 May 2026
Safety

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning

DGX agent

arXiv:2605.22511v1 Announce Type: cross Abstract: Post-training has become the dominant recipe for turning a language model into a competent search-augmented reasoning agent. A line of recent work pus

safetyarxiv-cs-cl
22 May 2026
Model Releases

Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs

DGX agent

arXiv:2605.22654v1 Announce Type: new Abstract: Previous detection studies have shown that LLMs cannot be effectively used as detectors, but these studies have not addressed modern Chinese poetry. Mor

model-releasesarxiv-cs-cl
22 May 2026
Safety

Self-Policy Distillation via Capability-Selective Subspace Projection

DGX agent

arXiv:2605.22675v1 Announce Type: new Abstract: Self-distillation bootstraps large language models (LLMs) by training on their own generations. However, existing methods either rely on external signal

safetyarxiv-cs-cl
22 May 2026
Research

Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews

DGX agent

arXiv:2605.21713v1 Announce Type: new Abstract: How can we distinguish whether a peer review was written by a human or generated by an AI model? We argue that, in this setting, authorship should not b

researcharxiv-cs-cl
22 May 2026
Model Releases

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

DGX agent

arXiv:2602.08064v2 Announce Type: replace-cross Abstract: The long-standing tension between Pre- and Post-Norm remains an open problem in Transformer architecture, reflecting a fundamental trade-off b

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation

DGX agent

arXiv:2605.22536v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in spatial intelligence, yet existing spatial reasoning benchmarks largely assume pr

model-releasesarxiv-cs-cl
22 May 2026
Agents

SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents

DGX agent

arXiv:2605.21965v1 Announce Type: new Abstract: Large language models increasingly use external tools such as web search and document retrieval to solve information-intensive tasks. However, multi-hop

agentsarxiv-cs-cl
22 May 2026
Agents

STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarking

DGX agent

arXiv:2507.03674v3 Announce Type: replace Abstract: Extracting structured information from scientific literature is critical for accelerating discovery, yet Large Language Models (LLMs) often struggle

agentsarxiv-cs-cl
22 May 2026
Model Releases

Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrieval

DGX agent

arXiv:2601.20107v2 Announce Type: replace-cross Abstract: Recent Vision-Language Models (e.g., ColPali) enable fine-grained Visual Document Retrieval (VDR) but incur prohibitive multi-vector index sto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

DGX agent

arXiv:2605.22202v1 Announce Type: new Abstract: In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding

model-releasesarxiv-cs-cl
22 May 2026
Local Ai

Structured-Sparse Attention for Entity Tracking with Subquadratic Sequence Complexity

DGX agent

arXiv:2605.22476v1 Announce Type: cross Abstract: Entity tracking requires maintaining and updating latent states for entities and attributes over long sequences. Recent task-specific attention operat

local-aiarxiv-cs-cl
22 May 2026
Model Releases

Sub-exponential Growth Dynamics in Complex Systems: A Piecewise Power-Law Model for the Diffusion of New Words and Names

DGX agent

arXiv:2511.04106v5 Announce Type: replace-cross Abstract: The diffusion of ideas and language in society has conventionally been described by S-shaped models, such as the logistic curve. However, the

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

DGX agent

arXiv:2605.22217v1 Announce Type: cross Abstract: Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent s

model-releasesarxiv-cs-cl
22 May 2026
Agents

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

DGX agent

arXiv:2605.22564v1 Announce Type: new Abstract: Today, tool-calling agents are commonly evaluated or tested on static datasets of execution traces, including input commands, agent responses, and assoc

agentsarxiv-cs-cl
22 May 2026
Model Releases

Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation

DGX agent

arXiv:2605.21491v1 Announce Type: cross Abstract: As language models accelerate scientific research by automating hypothesis generation and implementation, a new bottleneck emerges: evaluating and fil

model-releasesarxiv-cs-cl
22 May 2026
Safety

The Double Dilemma in Multi-Task Radiology Report Generation: A Gradient Dynamics Analysis and Solution

DGX agent

arXiv:2605.22635v1 Announce Type: cross Abstract: While multi-task learning based automatic radiology report generation (RRG) is widely adopted to ensure clinical consistency, most focus on architectu

safetyarxiv-cs-cl
22 May 2026
Model Releases

Token-Level LLM Collaboration via FusionRoute

DGX agent

arXiv:2601.05106v4 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a singl

model-releasesarxiv-cs-cl
22 May 2026
← Previous
1…8283848586…162
Next →