AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
22 May 2026

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

Model ReleasesDGX agent

arXiv:2605.21625v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus

FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing

SafetyDGX agent

arXiv:2605.22057v1 Announce Type: new Abstract: Enterprise routers assign queries to expert agents, yet deployed profiles stay static while agents evolve (prompts, tools, models), and developers rarel

From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.22462v1 Announce Type: new Abstract: We propose a five-stage methodology for causal feature analysis in transformer language models (probe design, feature extraction, causal validation, rob

From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

Model ReleasesDGX agent

arXiv:2605.21558v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts hav

From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning

ResearchDGX agent

arXiv:2605.22074v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but outcome-based RLVR remains inefficient on hard p

From TF-IDF to Transformers: A Comparative and Ensemble Approach to Sentiment Classification

ResearchDGX agent

arXiv:2605.22003v1 Announce Type: new Abstract: Sentiment analysis, also referred to as opinion mining, primarily tries to extract opinion from any text-based data. In the context of movie reviews and

General Agentic Planning Through Simulative Reasoning with World Models

AgentsDGX agent

arXiv:2507.23773v3 Announce Type: replace-cross Abstract: What does it mean to plan? Current agentic systems, whether scaffolded workflows or end-to-end policies, rely on reactive decision-making: sel

Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift

ResearchDGX agent

arXiv:2605.21849v1 Announce Type: cross Abstract: Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers s

GHI: Graphormer over Conditioned Hypergraph Incidence for Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.22228v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) requires models to bind sentiment evidence to the correct aspect, making it a natural testbed for fine-grained st

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

Model ReleasesDGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting

SafetyDGX agent

arXiv:2605.22258v1 Announce Type: new Abstract: Large language models (LLMs) require robust toxicity evaluation beyond explicit wording. This setting remains underexplored in Chinese, where toxicity m

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

Model ReleasesDGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

Model ReleasesDGX agent

arXiv:2605.22064v1 Announce Type: new Abstract: Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B,

HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.22035v1 Announce Type: cross Abstract: Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge

Hypergraph as Language

Local AiDGX agent

arXiv:2605.21858v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally g

IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

Model ReleasesDGX agent

arXiv:2605.22247v1 Announce Type: new Abstract: Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, th

ImProver: Agent-Based Automated Proof Optimization

AgentsDGX agent

arXiv:2410.04753v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been used to generate formal proofs of mathematical theorems in proofs assistants such as Lean. However, we

In Silico Modeling of the RAMPHO Buffer: Dissociating Informational and Energetic Masking via Phonetic Entropy in Deep Neural Networks

ResearchDGX agent

arXiv:2605.22465v1 Announce Type: new Abstract: The fundamental challenge of listening in multi-talker environments is a cognitive bottleneck, defined by the Ease of Language Understanding (ELU) model

InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models

Model ReleasesDGX agent

arXiv:2602.23200v2 Announce Type: replace-cross Abstract: When transformer-based language models are deployed for text generation, most of the inference time is spent in the decoding stage, where outp

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

Local AiDGX agent

arXiv:2511.07885v4 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strains t

Internal narratives parameterise affective states

ResearchDGX agent

arXiv:2502.09487v3 Announce Type: replace Abstract: Characterising how we verbalise our feelings is central to psychological assessment and intervention, yet the mapping between narrative and affectiv

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

Model ReleasesDGX agent

arXiv:2605.22079v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to generate structured outputs such as JSON, SQL, and code, yet public resources remain limited for evaluat

La representacion de la variacion contextual mediante definiciones terminologicas flexibles

TutorialsDGX agent

arXiv:1607.06330v2 Announce Type: replace Abstract: In this doctoral thesis, we apply premises of cognitive linguistics to terminological definitions and present a proposal called the flexible termino

LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance

SafetyDGX agent

arXiv:2605.22567v1 Announce Type: new Abstract: Reinforcement learning has proven effective for enhancing multi-step reasoning in large language models (LLMs), yet its benefits have not fully translat

LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

ResearchDGX agent

arXiv:2605.22012v1 Announce Type: new Abstract: Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasonin

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?

ResearchDGX agent

arXiv:2510.07962v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated remarkable progress in reasoning, often through supervised fine-tuning (SFT). However, SFT is resourc

Linear Dynamics in the RLVR Training of Large Language Models

Model ReleasesDGX agent

arXiv:2601.04537v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven significant performance gains in reasoning-oriented large language models (LL

LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications

Model ReleasesDGX agent

arXiv:2603.27355v2 Announce Type: replace-cross Abstract: We present a readiness harness for LLM and RAG applications that turns evaluation into a deployment decision workflow. The system combines aut

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

Model ReleasesDGX agent

arXiv:2605.22177v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing f

MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models

Model ReleasesDGX agent

arXiv:2510.23090v2 Announce Type: replace Abstract: Recent advances have investigated the use of pretrained large language models (LLMs) for time-series forecasting by aligning numerical inputs with L

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

Model ReleasesDGX agent

arXiv:2605.21796v1 Announce Type: cross Abstract: Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision

Modeling Pathology-Like Behavioral Patterns in Language Models Through Behavioral Fine-Tuning

SafetyDGX agent

arXiv:2605.22356v1 Announce Type: new Abstract: Large language models are increasingly used as computational tools for modeling human-like behavior. We introduce a behavioral induction framework that

Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora

SafetyDGX agent

arXiv:2605.22660v1 Announce Type: new Abstract: Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slang, and cultur

More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts

ResearchDGX agent

arXiv:2605.22641v1 Announce Type: new Abstract: Detecting Schwartz values in political text is difficult because implicit cues often depend on surrounding arguments and fine-grained distinctions betwe

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

Multi-Stage Training for Abusive Comment Detection in Indic Languages

ResearchDGX agent

arXiv:2605.22380v1 Announce Type: new Abstract: In recent years social media has become an increasingly popular tool for communication. People use it to share their ideas, exchange information, and di

NaviAgent: Graph-Driven Bilevel Planning for Scalable Tool Orchestration

SafetyDGX agent

arXiv:2506.19500v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) increasingly act as function-call agents that invoke external tools to tackle tasks beyond their static knowledge

One prompt is not enough: Instruction Sensitivity Undermines Embedding Model Evaluation

ResearchDGX agent

arXiv:2605.22544v1 Announce Type: new Abstract: Instruction embedding models have become common among state-of-the-art models, however are evaluated using a single prompt per task. The single-point ev

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI

SafetyDGX agent

arXiv:2507.05660v3 Announce Type: replace-cross Abstract: Customizing Large Language Models (LLMs) on untrusted datasets poses severe risks of injecting toxic behaviors. In this work, we introduce Opt

Pattern-and-root inflectional morphology: the Arabic broken plural

ResearchDGX agent

arXiv:2605.22310v1 Announce Type: new Abstract: We present a substantially implemented model of description of the inflectional morphology of Arabic nouns, with special attention to the management of

Planning in the LLM Era: Building for Reliability and Efficiency

ResearchDGX agent

arXiv:2605.21902v1 Announce Type: cross Abstract: Growing attention to intelligent agents has put a spotlight on one of their central capabilities: planning. Early attempts to leverage large language

Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation

Model ReleasesDGX agent

arXiv:2605.22487v1 Announce Type: new Abstract: Recent advances in Multilingual Large Language Models (MLLMs) have significantly enhanced cross-lingual conversational capabilities, yet modeling cultur

Probabilistic Attribution For Large Language Models

ResearchDGX agent

arXiv:2605.21726v1 Announce Type: new Abstract: The generative nature of Large Language Models (LLMs) is reflected in the conditional probabilities they compute to sample each response token given the

PromptNCE: Pointwise Mutual Information Predictions Using Only LLMs and Contrastive Estimation Prompts

Model ReleasesDGX agent

arXiv:2605.21776v1 Announce Type: new Abstract: Estimating mutual information from text usually requires training a task-specific critic, which limits its use in low-data settings. We ask whether larg

Psy-Chronicle:A Structured Pipeline for Synthesizing Long-Horizon Campus Psychological Counseling Dialogues

AgentsDGX agent

arXiv:2605.22140v1 Announce Type: new Abstract: In recent years, large language models have shown substantial potential in psychological support tasks. However, existing psychological counseling data

Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP

Model ReleasesDGX agent

arXiv:2603.20405v2 Announce Type: replace-cross Abstract: We report on an experiment in which Claude Opus~4.6, equipped with a suite of Model Context Protocol (MCP) tools for the Rocq proof assistant,

Quantizing Whisper-small: How design choices affect ASR performance

ApplicationsDGX agent

arXiv:2511.08093v2 Announce Type: replace-cross Abstract: Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computa

RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems

Model ReleasesDGX agent

arXiv:2510.13910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallu

RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator

Model ReleasesDGX agent

arXiv:2605.21748v1 Announce Type: new Abstract: As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes.

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

Model ReleasesDGX agent

arXiv:2605.22148v1 Announce Type: cross Abstract: Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent evaluation

Reducing Political Manipulation with Consistency Training

SafetyDGX agent

arXiv:2605.22771v1 Announce Type: new Abstract: Large language models (LLMs) exhibit systematic political bias across a variety of sensitive contexts. We find that LLMs handle counterpart topics from

Reflecti-Mate: A Conversational Agent for Adaptive Decision-Making Support Through System 1 and System 2 Thinking

AgentsDGX agent

arXiv:2605.22509v1 Announce Type: cross Abstract: Making high-stakes personal decisions involves cognitive, emotional, and intuitive processes, and individuals differ in how they allocate attention ac

Reflective Prompt Tuning through Language Model Function-Calling

Model ReleasesDGX agent

arXiv:2605.21781v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly capable of following instructions and complex reasoning, making prompting a flexible interface for

Residual Skill Optimization for Text-to-SQL Ensembles

Model ReleasesDGX agent

arXiv:2605.21792v1 Announce Type: new Abstract: Text-to-SQL ensembles improve over single-candidate generation by drawing multiple SQL candidates and selecting one, but their effectiveness is bounded

Scene Abstraction for Lexical Semantics: Structured Representations of Situated Meaning

ResearchDGX agent

arXiv:2605.22542v1 Announce Type: new Abstract: Coffee and tea share many properties, yet they evoke strikingly different situations, atmospheres, and affective associations. These situated dimensions

Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning

SafetyDGX agent

arXiv:2605.22511v1 Announce Type: cross Abstract: Post-training has become the dominant recipe for turning a language model into a competent search-augmented reasoning agent. A line of recent work pus

Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs

Model ReleasesDGX agent

arXiv:2605.22654v1 Announce Type: new Abstract: Previous detection studies have shown that LLMs cannot be effectively used as detectors, but these studies have not addressed modern Chinese poetry. Mor

Self-Policy Distillation via Capability-Selective Subspace Projection

SafetyDGX agent

arXiv:2605.22675v1 Announce Type: new Abstract: Self-distillation bootstraps large language models (LLMs) by training on their own generations. However, existing methods either rely on external signal

Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews

ResearchDGX agent

arXiv:2605.21713v1 Announce Type: new Abstract: How can we distinguish whether a peer review was written by a human or generated by an AI model? We argue that, in this setting, authorship should not b

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

Model ReleasesDGX agent

arXiv:2602.08064v2 Announce Type: replace-cross Abstract: The long-standing tension between Pre- and Post-Norm remains an open problem in Transformer architecture, reflecting a fundamental trade-off b

← Previous
1…6566676869…130
Next →