AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
23 Apr 2026

DRIV-EX: Counterfactual Explanations for Driving LLMs

SafetyDGX agent

arXiv:2603.00696v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as reasoning engines in autonomous driving, yet their decision-making remains opaque. We propose

Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving

AgentsDGX agent

arXiv:2604.20183v1 Announce Type: new Abstract: Large Language Models (LLMs) often struggle with structural ambiguity in optimization problems, where a single problem admits multiple related but confl

Duluth at SemEval-2026 Task 6: DeBERTa with LLM-Augmented Data for Unmasking Political Question Evasions

Model ReleasesDGX agent

arXiv:2604.20168v1 Announce Type: new Abstract: This paper presents the Duluth approach to SemEval-2026 Task 6 on CLARITY: Unmasking Political Question Evasions. We address Task 1 (clarity-level class


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Effects of Cross-lingual Evidence in Multilingual Medical Question Answering

ResearchDGX agent

arXiv:2604.20531v1 Announce Type: new Abstract: This paper investigates Multilingual Medical Question Answering across high-resource (English, Spanish, French, Italian) and low-resource (Basque, Kazak

ESGLens: An LLM-Based RAG Framework for Interactive ESG Report Analysis and Score Prediction

ResearchDGX agent

arXiv:2604.19779v1 Announce Type: new Abstract: Environmental, Social, and Governance (ESG) reports are central to investment decision-making, yet their length, heterogeneous content, and lack of stan

Evidence of Layered Positional and Directional Constraints in the Voynich Manuscript: Implications for Cipher-Like Structure

Model ReleasesDGX agent

arXiv:2604.19762v1 Announce Type: new Abstract: The Voynich Manuscript (VMS) exhibits a script of uncertain origin whose grapheme sequences have resisted linguistic analysis. We present a systematic a

Finding Duplicates in 1.1M BDD Steps: cukereuse, a Paraphrase-Robust Static Detector for Cucumber and Gherkin

Model ReleasesDGX agent

arXiv:2604.20462v1 Announce Type: cross Abstract: Behaviour-Driven Development (BDD) suites accumulate step-text duplication whose maintenance cost is established in prior work. Existing detection tec

Foundational Design Principles and Patterns for Building Robust and Adaptive GenAI-Native Systems

TutorialsDGX agent

arXiv:2508.15411v3 Announce Type: replace-cross Abstract: Generative AI (GenAI) has emerged as a transformative technology, demonstrating remarkable capabilities across diverse application domains. Ho

From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents

Model ReleasesDGX agent

arXiv:2604.20006v1 Announce Type: new Abstract: Personalized agents that interact with users over long periods must maintain persistent memory across sessions and update it as circumstances change. Ho

Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs

SafetyDGX agent

arXiv:2604.20382v1 Announce Type: new Abstract: Rising demand for mental health support has increased interest in using Large Language Models (LLMs) for counseling. However, adapting LLMs to this high

HaS: Accelerating RAG through Homology-Aware Speculative Retrieval

AgentsDGX agent

arXiv:2604.20452v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) expands the knowledge boundary of large language models (LLMs) at inference by retrieving external documents as c

How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues

Model ReleasesDGX agent

arXiv:2604.19783v1 Announce Type: new Abstract: Which persuasion strategies, if any, are associated with donation compliance? Answering this requires fine-grained strategy labels across a full corpus

How to measure the optimality of word or gesture order with respect to the principle of swap distance minimization

TutorialsDGX agent

arXiv:2604.01938v3 Announce Type: replace Abstract: The structure of all the permutations of a sequence can be represented as a permutohedron, a graph where vertices are permutations and two vertices

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models

ResearchDGX agent

arXiv:2604.19786v1 Announce Type: new Abstract: Evaluating humor in large language models (LLMs) is an open challenge because existing approaches yield isolated, incomparable metrics rather than unifi

Hybrid Multi-Phase Page Matching and Multi-Layer Diff Detection for Japanese Building Permit Document Review

Model ReleasesDGX agent

arXiv:2604.19770v1 Announce Type: new Abstract: We present a hybrid multi-phase page matching algorithm for automated comparison of Japanese building permit document sets. Building permit review in Ja

Improving End-to-End Training of Retrieval-Augmented Generation Models via Joint Stochastic Approximation

ResearchDGX agent

arXiv:2508.18168v3 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) has become a widely recognized paradigm to combine parametric memory with non-parametric memories. An RAG model

Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments

ApplicationsDGX agent

arXiv:2511.17069v3 Announce Type: replace Abstract: AI-driven automated scoring systems offer scalable and efficient means of evaluating complex student-generated responses. Yet, despite increasing de

Intersectional Fairness in Large Language Models

Model ReleasesDGX agent

arXiv:2604.20677v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in socially sensitive settings, raising concerns about fairness and biases, particularly across i

Knapsack Optimization-based Schema Linking for LLM-based Text-to-SQL Generation

Model ReleasesDGX agent

arXiv:2502.12911v3 Announce Type: replace Abstract: Generating SQLs from user queries is a long-standing challenge, where the accuracy of initial schema linking significantly impacts subsequent SQL ge

Language-Coupled Reinforcement Learning for Multilingual Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2601.14896v2 Announce Type: replace Abstract: Multilingual retrieval-augmented generation (MRAG) requires models to effectively acquire and integrate beneficial external knowledge from multiling

Large language models perceive cities through a culturally uneven baseline

SafetyDGX agent

arXiv:2604.20048v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to describe, evaluate and interpret places, yet it remains unclear whether they do so from a cultural

Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework

Model ReleasesDGX agent

arXiv:2604.20090v1 Announce Type: new Abstract: Cross-lingual chain-of-thought (XCoT) with self-consistency markedly enhances multilingual reasoning, yet existing methods remain costly due to extensiv

LLAMADRS: Evaluating Open-Source LLMs on Real Clinical Interviews--To Reason or Not to Reason?

Model ReleasesDGX agent

arXiv:2501.03624v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored.

LLM StructCore: Schema-Guided Reasoning Condensation and Deterministic Compilation

ResearchDGX agent

arXiv:2604.20560v1 Announce Type: new Abstract: Automatically filling Case Report Forms (CRFs) from clinical notes is challenging due to noisy language, strict output contracts, and the high cost of f

LoRA-FA: Efficient and Effective Low Rank Representation Fine-tuning

Model ReleasesDGX agent

arXiv:2308.03303v2 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) is crucial for improving their performance on downstream tasks, but full-parameter fine-tuning (Full-FT) is

Markov reads Pushkin, again: A statistical journey into the poetic world of Evgenij Onegin

ResearchDGX agent

arXiv:2604.20221v1 Announce Type: new Abstract: This study applies symbolic time series analysis and Markov modeling to explore the phonological structure of Evgenij Onegin-as captured through a graph

Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy

ResearchDGX agent

arXiv:2601.02989v2 Announce Type: replace Abstract: Large language models (LLMs), despite strong performance on complex mathematical problems, exhibit systematic limitations in counting tasks. This is

Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMs

SafetyDGX agent

arXiv:2601.02931v2 Announce Type: replace Abstract: Autoregressive LLMs perform well on relational tasks that require linking entities via relational words (e.g., father/son, friend), but it is unclea

MOA: Multi-Objective Alignment for Role-Playing Agents

SafetyDGX agent

arXiv:2512.09756v2 Announce Type: replace Abstract: Role-playing agents (RPAs) require balancing multiple objectives, such as instruction following, persona consistency, and stylistic fidelity, which

Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models

ResearchDGX agent

arXiv:2506.02132v5 Announce Type: replace Abstract: Large transformer-based language models dominate modern NLP, yet our understanding of how they encode linguistic information relies primarily on stu

Multi-Perspective Evidence Synthesis and Reasoning for Unsupervised Multimodal Entity Linking

ResearchDGX agent

arXiv:2604.20283v1 Announce Type: new Abstract: Multimodal Entity Linking (MEL) is a fundamental task in data management that maps ambiguous mentions with diverse modalities to the multimodal entities

Neural Bandit Based Optimal LLM Selection for a Pipeline of Subtasks

AgentsDGX agent

arXiv:2508.09958v3 Announce Type: replace Abstract: As large language models (LLMs) become increasingly popular, there is a growing need to predict which out of a set of LLMs will yield a successful a

'Newspaper Eat' Means 'Not Tasty': A Taxonomy and Benchmark for Coded Language in Real-World Chinese Online Reviews

Model ReleasesDGX agent

arXiv:2601.19932v2 Announce Type: replace Abstract: Coded language is an important part of human communication. It refers to cases where users intentionally encode meaning so that the surface text dif

Not all ANIMALs are equal: metaphorical framing through source domains and semantic frames

ResearchDGX agent

arXiv:2604.20454v1 Announce Type: new Abstract: Metaphors are powerful framing devices, yet their source domains alone do not fully explain the specific associations they evoke. We argue that the inte

On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks

ResearchDGX agent

arXiv:2604.20079v1 Announce Type: cross Abstract: Auto-regressive Large Language Models (LLMs) achieve strong performance on coding tasks, but incur high memory and inference costs. Diffusion-based la

Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization

ResearchDGX agent

arXiv:2601.12078v2 Announce Type: replace Abstract: Large language models (LLMs) excel at general-purpose tasks, yet adapting their responses to individual users remains challenging. Retrieval augment

Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs

ResearchDGX agent

arXiv:2603.27518v2 Announce Type: replace Abstract: Aligned language models that are trained to refuse harmful requests also exhibit over-refusal: they decline safe instructions that seemingly resembl

Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL

Model ReleasesDGX agent

arXiv:2604.20835v1 Announce Type: new Abstract: Modern language models demonstrate impressive coding capabilities in common programming languages (PLs), such as C++ and Python, but their performance i

PLR: Plackett-Luce for Reordering In-Context Learning Examples

Model ReleasesDGX agent

arXiv:2603.21373v2 Announce Type: replace-cross Abstract: In-context learning (ICL) adapts large language models by conditioning on a small set of ICL examples, avoiding costly parameter updates. Amon

RADS: Reinforcement Learning-Based Sample Selection Improves Transfer Learning in Low-resource and Imbalanced Clinical Settings

ApplicationsDGX agent

arXiv:2604.20256v1 Announce Type: new Abstract: A common strategy in transfer learning is few shot fine-tuning, but its success is highly dependent on the quality of samples selected as training examp

RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering

Model ReleasesDGX agent

arXiv:2604.20738v1 Announce Type: new Abstract: We introduce a benchmark dataset for question answering and translation in bilingual Latin and English settings, containing about 7,800 question-answer

Rethinking Reinforcement Fine-Tuning in LVLM: Convergence, Reward Decomposition, and Generalization

SafetyDGX agent

arXiv:2604.19857v1 Announce Type: cross Abstract: Reinforcement fine-tuning with verifiable rewards (RLVR) has emerged as a powerful paradigm for equipping large vision-language models (LVLMs) with ag

Retrofitting Small Multilingual Models for Retrieval: Matching 7B Performance with 300M Parameters

ResearchDGX agent

arXiv:2510.14274v2 Announce Type: replace Abstract: Training effective multilingual embedding models presents unique challenges due to the diversity of languages and task objectives. Although small mu

RExBench: Can coding agents autonomously implement AI research extensions?

Model ReleasesDGX agent

arXiv:2506.22598v3 Announce Type: replace Abstract: Agents based on Large Language Models (LLMs) have shown promise for performing sophisticated software engineering tasks autonomously. In addition, t

SAKE: Self-aware Knowledge Exploitation-Exploration for Grounded Multimodal Named Entity Recognition

Local AiDGX agent

arXiv:2604.20146v1 Announce Type: cross Abstract: Grounded Multimodal Named Entity Recognition (GMNER) aims to extract named entities and localize their visual regions within image-text pairs, serving

Self-Aware Vector Embeddings for Retrieval-Augmented Generation: A Neuroscience-Inspired Framework for Temporal, Confidence-Weighted, and Relational Knowledge

Model ReleasesDGX agent

arXiv:2604.20598v1 Announce Type: cross Abstract: Modern retrieval-augmented generation (RAG) systems treat vector embeddings as static, context-free artifacts: an embedding has no notion of when it w

SignDATA: Data Pipeline for Sign Language Translation

SafetyDGX agent

arXiv:2604.20357v1 Announce Type: cross Abstract: Sign-language datasets are difficult to preprocess consistently because they vary in annotation schema, clip timing, signer framing, and privacy const

SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks

Model ReleasesDGX agent

arXiv:2604.20087v1 Announce Type: new Abstract: Skills have become the de facto way to enable LLM agents to perform complex real-world tasks with customized instructions, workflows, and tools, but how

Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference

ApplicationsDGX agent

arXiv:2604.19943v1 Announce Type: new Abstract: Annotation pipelines in Natural Language Processing (NLP) commonly assume a single latent ground truth per instance and resolve disagreement through lab

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework

ResearchDGX agent

arXiv:2509.21267v3 Announce Type: replace Abstract: Large language models often generate homogeneous outputs, but whether this is problematic depends on the specific task. For objective math tasks, re

Text-to-Distribution Prediction with Quantile Tokens and Neighbor Context

Model ReleasesDGX agent

arXiv:2604.20216v1 Announce Type: new Abstract: Many applications of LLM-based text regression require predicting a full conditional distribution rather than a single point value. We study distributio

The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models

Model ReleasesDGX agent

arXiv:2604.20225v1 Announce Type: new Abstract: Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current bench

The Imperfective Paradox in Large Language Models

SafetyDGX agent

arXiv:2601.09373v2 Announce Type: replace Abstract: Do Large Language Models (LLMs) genuinely grasp the compositional semantics of events, or do they rely on surface-level probabilistic heuristics? We

To Know is to Construct: Schema-Constrained Generation for Agent Memory

Model ReleasesDGX agent

arXiv:2604.20117v1 Announce Type: new Abstract: Constructivist epistemology argues that knowledge is actively constructed rather than passively copied. Despite the generative nature of Large Language

Towards High-Quality Machine Translation for Kokborok: A Low-Resource Tibeto-Burman Language of Northeast India

Model ReleasesDGX agent

arXiv:2604.19778v1 Announce Type: new Abstract: We present KokborokMT, a high-quality neural machine translation (NMT) system for Kokborok (ISO 639-3), a Tibeto-Burman language spoken primarily in Tri

Tracing Relational Knowledge Recall in Large Language Models

ResearchDGX agent

arXiv:2604.19934v1 Announce Type: new Abstract: We study how large language models recall relational knowledge during text generation, with a focus on identifying latent representations suitable for r

Trajectory2Task: Training Robust Tool-Calling Agents with Synthesized Yet Verifiable Data for Complex User Intents

Model ReleasesDGX agent

arXiv:2601.20144v3 Announce Type: replace Abstract: Tool-calling agents are increasingly deployed in real-world customer-facing workflows. Yet most studies on tool-calling agents focus on idealized se

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.20398v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at function-level code generation, project-level tasks such as generating functional and visually aesthetic mul

What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization

ApplicationsDGX agent

arXiv:2601.17609v2 Announce Type: replace Abstract: In domains like medicine and finance, large-scale labeled data is costly and often unavailable, leading to models trained on small datasets that str

Where Reasoning Breaks: Logic-Aware Path Selection by Controlling Logical Connectives in LLMs Reasoning Chains

Local AiDGX agent

arXiv:2604.20564v1 Announce Type: new Abstract: While LLMs demonstrate impressive reasoning capabilities, they remain fragile in multi-step logical deduction, where a single transition error can propa

← Previous
1…101102103104105…129
Next →