AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
25 May 2026

SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval

ResearchDGX agent

arXiv:2601.03260v2 Announce Type: replace-cross Abstract: AI agents have seen widespread adoption in information retrieval for scientific research, giving rise to tools such as Deep Research. However,

Self-Improving In-Context Learning

ResearchDGX agent

arXiv:2605.23180v1 Announce Type: new Abstract: We propose to improve in-context learning (ICL) by optimizing the continuous embeddings of a fixed few-shot prompt at test time. The key observation is

SemEval-2026 Task 6: CLARITY -- Unmasking Political Question Evasions

Model ReleasesDGX agent

arXiv:2603.14027v2 Announce Type: replace Abstract: Political speakers often avoid answering questions directly while maintaining the appearance of responsiveness. Despite its importance for public di


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation

Model ReleasesDGX agent

arXiv:2412.14642v4 Announce Type: replace Abstract: Recently, Large Language Models (LLMs) have demonstrated great potential in natural language-driven molecule discovery. However, existing datasets a

Strong Teacher Not Needed? On Distillation in LLM Pretraining

ResearchDGX agent

arXiv:2605.23857v1 Announce Type: cross Abstract: Knowledge distillation generally assumes a strong-to-weak relationship where stronger teachers yield better students. In this work, we examine this as

Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts

ApplicationsDGX agent

arXiv:2605.23597v1 Announce Type: new Abstract: Matching person names across heterogeneous records is a core challenge in entity resolution, especially within linguistically and culturally complex env

TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration

Model ReleasesDGX agent

arXiv:2602.08404v2 Announce Type: replace Abstract: Diffusion large language models (dLLMs) have recently gained significant attention due to their inherent support for parallel decoding. Building on

The Efficiency Frontier: A Unified Framework for Cost-Performance Optimization in LLM Context Management

ResearchDGX agent

arXiv:2605.23071v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on long-context processing, but expanding context windows introduces substantial computational and financ

TurkicNLP: An NLP Toolkit for Turkic Languages

ResearchDGX agent

arXiv:2602.19174v5 Announce Type: replace Abstract: Natural language processing for the Turkic language family, spoken by over 200 million people across Eurasia, remains fragmented, with most language

Vector Retrieval with Similarity and Diversity: How Hard Is It?

Model ReleasesDGX agent

arXiv:2407.04573v4 Announce Type: replace-cross Abstract: Dense vector retrieval is an important building block of modern machine learning systems, underlying applications ranging from semantic search

What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference

ResearchDGX agent

arXiv:2605.23158v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) on resource-constrained devices remains challenging, spurring interest in split inference, where models

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA

Model ReleasesDGX agent

arXiv:2605.23067v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a viable recipe for training LLM agents to reason over external memory banks in multi-session dialogue. Exist

When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance

ApplicationsDGX agent

arXiv:2605.22975v1 Announce Type: new Abstract: We ask whether large language models (LLMs) treat queries about religious conversion symmetrically. The answer is no. When asked for advice on hypotheti

When Is Next-Token Prediction Useful? Marginalization, Ergodicity, Mixture Identifiability, Local Sufficiency, RAG, Tools, and Programming

AgentsDGX agent

arXiv:2605.23278v1 Announce Type: new Abstract: Language models trained on observed sequences are often described as learning the conditional distribution of the next token given previous tokens. This

When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening

Model ReleasesDGX agent

arXiv:2605.23148v1 Announce Type: new Abstract: As demand for mental health care outpaces clinician-delivered assessment, scalable screening tools are increasingly needed. Large language models (LLMs)

22 May 2026

A Comparative Study of Language Models for Khmer Retrieval-Augmented Question Answering

Model ReleasesDGX agent

arXiv:2605.22099v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for grounding large language model (LLM) outputs in retrieved evidence, thereby

A Tutorial on Diffusion Theory: From Differential Equations to Diffusion Models

TutorialsDGX agent

arXiv:2605.22586v1 Announce Type: cross Abstract: This tutorial develops diffusion models from the viewpoint of differential equations. We begin with the conditional Gaussian forward process and show

ACC: Compiling Agent Trajectories for Long-Context Training

AgentsDGX agent

arXiv:2605.21850v1 Announce Type: new Abstract: Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requires costly lo

Accelerated Test-Time Scaling with Model-Free Speculative Sampling

ResearchDGX agent

arXiv:2506.04708v3 Announce Type: replace Abstract: Language models have demonstrated remarkable capabilities in reasoning tasks through test-time scaling techniques like best-of-N sampling and tree s

Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

SafetyDGX agent

arXiv:2605.22608v1 Announce Type: new Abstract: Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious

AMEL: Accumulated Message Effects on LLM Judgments

Model ReleasesDGX agent

arXiv:2605.22714v1 Announce Type: cross Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing th

Amplifying, Not Learning: Fine-Tuned AI Text Detectors Amplify a Pretrained Direction

SafetyDGX agent

arXiv:2605.21653v1 Announce Type: cross Abstract: AI text detectors amplify a pretrained typicality axis; they do not construct an AI-vs-human boundary. On raw encoders before any task supervision, pr

An Entity Linking Agent for Question Answering

AgentsDGX agent

arXiv:2508.03865v4 Announce Type: replace Abstract: Some Question Answering (QA) systems rely on knowledge bases (KBs) to provide accurate answers. Entity Linking (EL) plays a critical role in linking

AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

TutorialsDGX agent

arXiv:2605.22715v1 Announce Type: cross Abstract: As wearable and mobile devices become increasingly embedded in daily life, they offer a practical way to continuously sense human motion in the wild.

ArabDiscrim: A Decade-Long Arabic Facebook Corpus on Racism and Discrimination

Model ReleasesDGX agent

arXiv:2605.22081v1 Announce Type: new Abstract: We present ArabDiscrim, a decade-long lexical resource and corpus of 293K public Arabic Facebook posts (2014--2024) discussing racism and discrimination

Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation

ResearchDGX agent

arXiv:2605.22435v1 Announce Type: new Abstract: Hate speech and misinformation frequently co-occur online, amplifying prejudice and polarization. Given their scale, using Large Language Models (LLMs)

Audience Engagement with Arabic Women's Social Empowerment and Wellbeing: A Decadal Corpus

Model ReleasesDGX agent

arXiv:2605.22204v1 Announce Type: new Abstract: This paper presents the Arabic Women and Society Corpus, a ten year collection of 252,487 public Arabic Facebook posts related to women's empowerment an

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model

Model ReleasesDGX agent

arXiv:2605.21728v1 Announce Type: cross Abstract: Image captioning evaluation remains a significant challenge, as vision-language models evolve toward more challenging capabilities such as generating

BeLink: Biomedical Entity Linking Meets Generative Re-Ranking

ApplicationsDGX agent

arXiv:2605.22501v1 Announce Type: new Abstract: Despite recent progress, Biomedical Entity Linking (BEL) with large language models (LLMs) remains computationally inefficient and challenging to deploy

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models

Model ReleasesDGX agent

arXiv:2605.22732v1 Announce Type: cross Abstract: We investigate whether acoustic emotion recognition models can serve as proxies for the Pathos dimension in political speech analysis, as operationali

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

Model ReleasesDGX agent

arXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)

Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansion

Model ReleasesDGX agent

arXiv:2605.22579v1 Announce Type: new Abstract: Recent work has identified a counterintuitive phenomenon termed 'Hyperfitting', where fine-tuning Large Language Models (LLMs) to near-zero training los

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

Model ReleasesDGX agent

arXiv:2605.22643v1 Announce Type: new Abstract: Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follo

Boundary-targeted Membership Inference Attacks on Safety Classifiers

Local AiDGX agent

arXiv:2605.22373v1 Announce Type: cross Abstract: Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with

Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries

Local AiDGX agent

arXiv:2605.21712v1 Announce Type: new Abstract: Transportation safety analysis requires integrating crash records, roadway attributes, and geospatial data through GIS-based workflows, but access remai

Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)

Model ReleasesDGX agent

arXiv:2605.22005v1 Announce Type: cross Abstract: We show that singular value decomposition of the lm_head} weight matrix of a transformer-based large language model -- requiring only five lines of Py

Chinese sensorimotor and embodiment norms for 3,000 lexicalized concepts

ResearchDGX agent

arXiv:2605.22616v1 Announce Type: new Abstract: Understanding how conceptual knowledge is grounded in bodily experience, and to what extent machine systems can acquire such knowledge without direct se

ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning

Model ReleasesDGX agent

arXiv:2605.22734v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) treat disease associations as static facts, but temporal information is crucial for clinical reasoning, e.g., a sympto

Claim-Selective Certification for High-Risk Medical Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2605.21949v1 Announce Type: new Abstract: Medical RAG systems in high-risk QA settings are often evaluated through a single answer-or-abstain decision, but mixed evidence may support one claim,

Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse

Model ReleasesDGX agent

arXiv:2605.22447v1 Announce Type: new Abstract: The study of online discourse has become central to understanding societal polarization. While much research has focused on detecting overt toxicity, th

Comparing LLM and Fine-Tuned Model Performance on NVDRS Circumstance Extraction with Varying Prompt Complexity

Model ReleasesDGX agent

arXiv:2605.21845v1 Announce Type: new Abstract: Suicide is a leading cause of death in the United States, and understanding the circumstances that precede it requires extracting structured information

CR4T: Rewrite-Based Guardrails for Adolescent LLM Safety

SafetyDGX agent

arXiv:2605.21609v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in adolescent digital environments, mediating information seeking, advice, and emotionally sensit

CritiSense: Critical Digital Literacy and Resilience Against Misinformation

ResearchDGX agent

arXiv:2603.16672v2 Announce Type: replace-cross Abstract: Misinformation on social media undermines informed decision-making and public trust. Prebunking offers a proactive complement by helping users

Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency

Model ReleasesDGX agent

arXiv:2605.22137v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate strong capabilities across various tasks, they exhibit significant performance discrepancies across la

DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QA

ResearchDGX agent

arXiv:2605.22411v1 Announce Type: new Abstract: Large language model (LLM) agents still struggle with long-term memory question answering, where answer-supporting evidence is often scattered across lo

Detecting Synthetic Political Narratives in Cross-Platform Social Media Discourse

ResearchDGX agent

arXiv:2605.21540v1 Announce Type: cross Abstract: The proliferation of large language models has introduced a new paradigm of synthetic political communication in which narratives may be generated, se

Diagnosis Is Not Prescription: Linguistic Co-Adaptation Explains Patching Hazards in LLM Pipelines

SafetyDGX agent

arXiv:2605.21958v1 Announce Type: new Abstract: When a multi-module LLM agent fails, the module most responsible for the failure is not necessarily the best place to intervene. We demonstrate this Dia

Discovering Implicit Large Language Model Alignment Objectives

SafetyDGX agent

arXiv:2602.15338v2 Announce Type: replace-cross Abstract: Large language model (LLM) alignment relies on complex reward signals that often obscure the specific behaviors being incentivized, creating c

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

ResearchDGX agent

arXiv:2605.22170v1 Announce Type: new Abstract: In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges abo

Does Slightly Mean Somewhat? Measuring Vague Intensity Words in LLM Numeric Actions

Model ReleasesDGX agent

arXiv:2605.21827v1 Announce Type: new Abstract: Do language models preserve the ordinal meaning of intensity words when those words must produce numeric actions? I study a researcher-constructed scale

Echo: Learning from Experience Data via User-Driven Refinement

AgentsDGX agent

arXiv:2605.21984v1 Announce Type: cross Abstract: Static 'human data' faces inherent limitations: it is expensive to scale and bounded by the knowledge of its creators. Continuous learning from 'exper

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

Model ReleasesDGX agent

arXiv:2605.22138v1 Announce Type: cross Abstract: How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thou

Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention

SafetyDGX agent

arXiv:2605.21842v1 Announce Type: cross Abstract: Standard transformer attention computes pairwise similarity between queries and keys, treating all tokens as equally salient regardless of their intri

Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning

ResearchDGX agent

arXiv:2401.00139v3 Announce Type: replace-cross Abstract: This paper introduces a causal attribution model to enhance the interpretability of large language models (LLMs) and improve their causal reas

EntmaxKV: Support-Aware Decoding for Entmax Attention

ResearchDGX agent

arXiv:2605.21649v1 Announce Type: cross Abstract: Long-context decoding is increasingly limited by KV-cache memory traffic since each generated token attends over a cache whose size grows linearly wit

Epicure: Navigating the Emergent Geometry of Food Ingredient Embeddings

ResearchDGX agent

arXiv:2605.22391v1 Announce Type: cross Abstract: We present Epicure, a family of three sibling skip-gram ingredient embeddings retrained from scratch on a multilingual recipe corpus. We aggregate 4.1

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark

Model ReleasesDGX agent

arXiv:2503.17599v3 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks pr

Evaluating Commercial AI Chatbots as News Intermediaries

Model ReleasesDGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

ResearchDGX agent

arXiv:2605.22203v1 Announce Type: new Abstract: In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Aug

← Previous
1…6364656667…129
Next →