AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
3 Jun 2026

Experience-Driven Dynamic Exits for LLMs with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.03113v1 Announce Type: new Abstract: Large Language Models suffer from slow autoregressive inference. While self-speculative decoding accelerates this process, its efficiency is hampered by

Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models

Local AiDGX agent

arXiv:2606.03780v1 Announce Type: new Abstract: Causal tracing of factual recall has been studied predominantly in dense transformer language models, where interventions localize information flow to l

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

SafetyDGX agent

arXiv:2606.03793v1 Announce Type: new Abstract: Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

FederatedSkill: Federated Learning for Agentic Skill Evolution

Local AiDGX agent

arXiv:2606.03143v1 Announce Type: cross Abstract: Modern LLM agents increasingly rely on skill libraries to handle complex tasks, making skill evolution a primary driver of self-improvement. However,

Framing Migration News with LLMs: Structured CoT as a Support for Human Interpretation

Local AiDGX agent

arXiv:2606.03761v1 Announce Type: new Abstract: Frame analysis of migration news is a socially consequential task: media scholars and researchers who study how migration is narrated need tools that ar

From Script to Semantics: Prompting Strategies for African NLI

Model ReleasesDGX agent

arXiv:2606.03304v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains

G^2C-MT: Graph-Guided Context Selection for Document-Level Machine Translation

Model ReleasesDGX agent

arXiv:2606.03078v1 Announce Type: new Abstract: Effective document-level machine translation (DocMT) requires capturing long-range discourse dependencies. Recent work has explored retrieval-based and

GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

SafetyDGX agent

arXiv:2606.03180v1 Announce Type: cross Abstract: Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workfl

Greener Than Humans? Environmental Attitudes in Large Language Models

Model ReleasesDGX agent

arXiv:2606.02741v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in sustainability-related decision support, reporting, and public communication, yet little systemati

Hallucination Is Linearly Decodable from Mid-Layer Hidden States in Quantized LLMs

Model ReleasesDGX agent

arXiv:2606.02628v1 Announce Type: cross Abstract: We investigate whether open-source LLMs encode a linearly separable truthfulness signal in their hidden states, and at which network depth this signal

Hint-Guided Diversified Policy Optimization for LLM Reasoning

SafetyDGX agent

arXiv:2606.03021v1 Announce Type: new Abstract: Recent developments in Large Language Models (LLMs) have showcased impressive reasoning capabilities, with Reinforcement Learning with Verifiable Reward

HybridThinker: Efficient Chain-of-Thought Reasoning via Compressed Memory and Transient Thought Steps

ResearchDGX agent

arXiv:2606.03768v1 Announce Type: new Abstract: Extended chain-of-thought (CoT) traces improve LLM reasoning but incur substantial computational and memory costs. While existing CoT compression method

HyperPatch: Sequential Knowledge Editing Under n-ary Structural Drift

Model ReleasesDGX agent

arXiv:2606.03179v1 Announce Type: new Abstract: Large Language Models (LLMs) rely on Knowledge Editing (KE) to maintain temporal validity, yet real-world knowledge is inherently n-ary. We demonstrate

Instant Personalized Large Language Model Adaptation via Hypernetwork

Model ReleasesDGX agent

arXiv:2510.16282v2 Announce Type: replace Abstract: Personalized large language models (LLMs) tailor content to individual preferences using user profiles or histories. However, existing parameter-eff

KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering

SafetyDGX agent

arXiv:2512.10999v3 Announce Type: replace Abstract: Knowledge Base Question Answering (KBQA) challenges models to bridge the gap between natural language and strict knowledge graph schemas by generati

KletterMix: Climbing Toward High-Quality German Pretraining Data

ResearchDGX agent

arXiv:2606.03773v1 Announce Type: new Abstract: High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their Engl

Knowledge Editing in Masked Diffusion Language Models

Model ReleasesDGX agent

arXiv:2606.03924v1 Announce Type: new Abstract: Knowledge editing aims to update or correct factual knowledge in a language model. A widely used approach, locate-then-edit, does this in two steps: it

Language Bias under Conflicting Information in Multilingual LLMs

Model ReleasesDGX agent

arXiv:2604.07123v2 Announce Type: replace Abstract: Large Language Models (LLMs) have been shown to contain biases in the process of integrating conflicting information when answering questions. Here

Language Models Compare Quantities Using Number-specific and Unit-specific Heuristics

ResearchDGX agent

arXiv:2606.03982v1 Announce Type: new Abstract: Quantities with measurement units, such as 110 cm and 1.2 m, require language models (LMs) to combine a numeral with a symbolic unit scale. Here, we stu

Large Language Models Are Overconfident in Their Own Responses

SafetyDGX agent

arXiv:2606.03437v1 Announce Type: new Abstract: Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, lit

Learning without training: The implicit dynamics of in-context learning

TutorialsDGX agent

arXiv:2507.16003v4 Announce Type: replace Abstract: One of the most striking features of Large Language Models (LLMs) is their ability to learn in-context. Namely at inference time an LLM is able to l

Letting Tutor Personas Speak Up for LLMs: Learning Steering Vectors from Dialogue via Preference Optimization

SafetyDGX agent

arXiv:2602.07639v2 Announce Type: replace Abstract: With the emergence of large language models (LLMs) as a powerful class of generative artificial intelligence (AI), their use in tutoring has become

Lexicons and grammars for language processing: industrial or handcrafted products?

ResearchDGX agent

arXiv:2606.03412v1 Announce Type: new Abstract: During the recent years, the use of linguistic data for language processing increased progressively. Such data are now commonly called language resource

Lingo_Research_Group at SemEval-2026 Task 9: Evaluating Prompt Variants for Polarization Detection

ResearchDGX agent

arXiv:2606.03334v1 Announce Type: new Abstract: Our submission presented in this paper is for SemEval-2026 Task 9: Multilingual Text Classification Challenge - Polarization Detection and it covers all

Linguistic Productivity in Large Language Models: Models Coerce, but do not Preempt

ResearchDGX agent

arXiv:2606.02953v1 Announce Type: new Abstract: Usage-based theories of grammars posit that creative productivity of the structures of language is both bolstered and constrained by two distinct freque

Memory Retrieval for Changing Preferences

ResearchDGX agent

arXiv:2606.02976v1 Announce Type: new Abstract: Long-context dialogue systems must decide both when to access memory and which parts of the interaction history are relevant. Existing approaches typica

MemTrain: Self-Supervised Context Memory Training

AgentsDGX agent

arXiv:2606.03197v1 Announce Type: new Abstract: Memory is an indispensable capability for long-horizon LLM agents, enabling them to preserve and utilize information accumulated across extended interac

Multi-component Causal Tracing in Large Language Models

SafetyDGX agent

arXiv:2606.03085v1 Announce Type: cross Abstract: Causal tracing systematically intervenes on a large language model's (LLM's) internal representations to uncover and quantify the causal pathways link

Multi-Segment Attention: Enabling Efficient KV-Cache Management for Faster Large Language Model Serving

SafetyDGX agent

arXiv:2606.02964v1 Announce Type: cross Abstract: Large Language Model (LLM) inference relies on key-value (KV) caches to avoid redundant attention computation. While approximate KV cache retention te

Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

Model ReleasesDGX agent

arXiv:2606.03291v1 Announce Type: new Abstract: Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. Howev

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models

ResearchDGX agent

arXiv:2602.03681v2 Announce Type: replace Abstract: The quadratic computational complexity of softmax transformers has become a bottleneck in long-context scenarios. In contrast, linear attention mode

Neuron Populations Exhibit Divergent Selectivity with Scale

ApplicationsDGX agent

arXiv:2606.03990v1 Announce Type: cross Abstract: We investigate whether neuron populations within neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables

On the Persistent Effects of Lexicality in Large Language Mod

ApplicationsDGX agent

arXiv:2606.02750v1 Announce Type: new Abstract: Representations extracted from large language models (LLMs) play an important role in many downstream applications. However, the structure of these repr

Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics

ResearchDGX agent

arXiv:2606.02981v1 Announce Type: new Abstract: Best-of-N inference scaling (drawing N candidate answers from a language model and returning the one a reward model ranks highest) improves accuracy by

Privacy-Aware Decoding: Mitigating Privacy Leakage of Large Language Models in Retrieval-Augmented Generation

ApplicationsDGX agent

arXiv:2508.03098v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) enhances the factual accuracy of large language models (LLMs) by conditioning outputs on external knowledge sou

PsychoPass: Geometric Profiling of Multi-Turn Adversarial LLM Conversations

SafetyDGX agent

arXiv:2606.03136v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks on large language models (LLMs) reveal a mismatch in current guardrails: they operate on individual turns, while attacks

Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains

ResearchDGX agent

arXiv:2505.16014v5 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems deployed in sensitive domains must provide interpretable evidence selection and robust safeguards again

Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA

Model ReleasesDGX agent

arXiv:2606.03728v1 Announce Type: new Abstract: Retrieval-augmented generation systems for legal question answering typically retrieve passages based on semantic similarity and provide them to a langu

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions

Model ReleasesDGX agent

arXiv:2606.03889v1 Announce Type: new Abstract: Agent benchmarks should reflect what users actually ask deployed agents to do, yet existing benchmarks often miss key realism properties of real develop

Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?

TutorialsDGX agent

arXiv:2606.03782v1 Announce Type: new Abstract: Large language models (LLMs) offer a promising approach to machine translation (MT) for extremely low-resource languages by incorporating linguistic res

Rethinking the Idiomaticity Decomposability Hypothesis: Evidence from Distributional Learning

ResearchDGX agent

arXiv:2606.03817v1 Announce Type: new Abstract: Idioms can be analysed in terms of their decomposability, the extent to which constituent meanings contribute to the figurative whole. Decomposability i

SagaQA: A Multi-hop Reasoning Benchmark for Long-form Narrative Understanding in TV Series

Model ReleasesDGX agent

arXiv:2606.03301v1 Announce Type: new Abstract: We introduce SagaQA, a long-form video benchmark for multi-hop reasoning over full-length TV series. Existing video reasoning benchmarks often emphasize

Sample-Size Scaling of the African Languages NLI Evaluation

Model ReleasesDGX agent

arXiv:2606.03219v1 Announce Type: new Abstract: African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performan

SEA-Embedding: Open and Reproducible Text Embeddings for Southeast Asia

ApplicationsDGX agent

arXiv:2606.03027v1 Announce Type: new Abstract: Text embeddings are fundamental to many downstream applications, making robustness important for real-world NLP. However, most recent state-of-the-art e

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

Model ReleasesDGX agent

arXiv:2606.03284v1 Announce Type: new Abstract: Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NL

See, Infer, Intervene: Proactive World Modeling for Goal-Oriented Social Intelligence

Model ReleasesDGX agent

arXiv:2606.03371v1 Announce Type: new Abstract: Multimodal retail agents should not only recognize what a customer is doing, but also decide whether and how to assist before an explicit request is mad

Selective Token-Level Cryptographic Redaction for Privacy-Preserving Clinical Deployment of Large Language Models

SafetyDGX agent

arXiv:2606.03399v1 Announce Type: new Abstract: While large language models (LLMs) are increasingly used for clinical applications, many existing pipelines require sending raw sensitive health informa

SenseJudge: Human-Centric Preference-Driven Judgment Framework

Model ReleasesDGX agent

arXiv:2606.03189v1 Announce Type: new Abstract: Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However

Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill

Model ReleasesDGX agent

arXiv:2606.03980v1 Announce Type: cross Abstract: Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) p

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling

ResearchDGX agent

arXiv:2606.03102v1 Announce Type: new Abstract: Test-time scaling improves the reasoning performance of large language models but incurs substantial cost in both total computation and latency. Existin

Social Caption: Evaluating Social Understanding in Multimodal Models

ResearchDGX agent

arXiv:2601.14569v2 Announce Type: replace Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions. We introduce SOCIAL

Speech Emotion Recognition using Attention-based LSTM-Network with Residual Connection

ResearchDGX agent

arXiv:2606.03359v1 Announce Type: cross Abstract: Speech emotion recognition is an important component of modern human-computer interaction systems. However, many state-of-the-art approaches rely on l

Structures Facilitate Retrieve, Rerank, and Generate

ResearchDGX agent

arXiv:2606.03247v1 Announce Type: new Abstract: Document-grounded dialogue systems (DGDS) utilize knowledge from external documents to answer domain-specific user questions. Existing solutions typical

The Deliberative Illusion: Diagnosing Factual Attrition and Stance Homogenization in Multi-Agent LLM Deliberation

AgentsDGX agent

arXiv:2606.03032v1 Announce Type: new Abstract: Multi-agent LLM systems often treat consensus as evidence of successful interaction. For deliberative problems, however, reliability depends on whether

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

Model ReleasesDGX agent

arXiv:2606.03043v1 Announce Type: new Abstract: LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared

The Ghost Annotator: a Framework to Explore Human Label Variation in Content Moderation through Conformal Prediction

SafetyDGX agent

arXiv:2606.02911v1 Announce Type: new Abstract: Current research primarily focuses on model performance, while comparatively less attention has been devoted to uncertainty estimation, particularly in

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

Model ReleasesDGX agent

arXiv:2606.03250v1 Announce Type: new Abstract: Digital healthcare generates vast amounts of clinical text that can support AI-assisted applications, yet German biomedical language models remain limit

Topics as Proxies for Sociodemographics: How Conversational Context Affects LLM Answers

ApplicationsDGX agent

arXiv:2606.02776v1 Announce Type: new Abstract: When large language models (LLMs) are used in high-stakes scenarios, such as legal, medical and financial advice, even a single conversation history is

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

SafetyDGX agent

arXiv:2509.22854v2 Announce Type: replace Abstract: Implicit in-context learning (ICL) has newly emerged as a promising paradigm that simulates ICL behaviors in the representation space of large langu

Translating Classical Poetry into Modern Prose

ResearchDGX agent

arXiv:2606.02806v1 Announce Type: new Abstract: We introduce Padyam2Gadyam, a dataset for the task of poem-to-prose translation from 13th-17th Century Telugu Classical Poetry to contemporary Telugu an

← Previous
1…4647484950…129
Next →