AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
17 Apr 2026

SeaAlert: Critical Information Extraction From Maritime Distress Communications with Large Language Models

SafetyDGX agent

arXiv:2604.14163v1 Announce Type: new Abstract: Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergencies at sea.

SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs

Local AiDGX agent

arXiv:2602.13529v2 Announce Type: replace-cross Abstract: Federated learning (FL) enables collaborative training across organizational silos without sharing raw data, making it attractive for privacy-

Segment-Level Coherence for Robust Harmful Intent Probing in LLMs

ResearchDGX agent

arXiv:2604.14865v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly exposed to adaptive jailbreaking, particularly in high-stakes Chemical, Biological, Radiological, and Nucl


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation

Model ReleasesDGX agent

arXiv:2604.14339v1 Announce Type: new Abstract: Large language models (LLMs) increasingly operate in settings that require reliable long-context understanding, such as retrieval-augmented generation a

Similarity-Distance-Magnitude Activations

ResearchDGX agent

arXiv:2509.12760v4 Announce Type: replace-cross Abstract: We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax

Social Story Frames: Contextual Reasoning about Narrative Intent and Reception

ResearchDGX agent

arXiv:2512.15925v2 Announce Type: replace Abstract: Reading stories evokes rich interpretive, affective, and evaluative responses, such as inferences about narrative intent or judgments about characte

SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models

SafetyDGX agent

arXiv:2604.14672v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used in urban planning, but since gendered space theory highlights how gender hierarchies are embedd

Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects

ApplicationsDGX agent

arXiv:2510.07890v3 Announce Type: replace Abstract: Research on cross-dialectal transfer from a standard to a non-standard dialect variety has typically focused on text data. However, dialects are pri

Stateful Evidence-Driven Retrieval-Augmented Generation with Iterative Reasoning

TutorialsDGX agent

arXiv:2604.14170v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) grounds Large Language Models (LLMs) in external knowledge but often suffers from flat context representations and

StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation

SafetyDGX agent

arXiv:2604.14631v1 Announce Type: new Abstract: Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan. Existing app

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models

ResearchDGX agent

arXiv:2512.23578v3 Announce Type: replace Abstract: In this paper, we show that when spoken language models (SLMs) are instructed to speak in a specific speaking style at the beginning of a multi-turn

Text2Arch: A Dataset for Generating Scientific Architecture Diagrams from Natural Language Descriptions

ApplicationsDGX agent

arXiv:2604.14941v1 Announce Type: new Abstract: Communicating complex system designs or scientific processes through text alone is inefficient and prone to ambiguity. A system that automatically gener

The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious

SafetyDGX agent

arXiv:2604.14414v1 Announce Type: new Abstract: Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models

Local AiDGX agent

arXiv:2604.14363v1 Announce Type: new Abstract: Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly understood.

The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows

SafetyDGX agent

arXiv:2604.14807v1 Announce Type: cross Abstract: The rapid integration of large language models (LLMs) into everyday workflows has transformed how individuals perform cognitive tasks such as writing,

The PICCO Framework for Large Language Model Prompting: A Taxonomy and Reference Architecture for Prompt Structure

SafetyDGX agent

arXiv:2604.14197v1 Announce Type: new Abstract: Large language model (LLM) performance depends heavily on prompt design, yet prompt construction is often described and applied inconsistently. Our purp

Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration

Model ReleasesDGX agent

arXiv:2507.02935v2 Announce Type: replace Abstract: Successful human-agent teaming relies on an agent being able to understand instructions given by a (human) principal. In many cases, an instruction

Three-Phase Transformer

Model ReleasesDGX agent

arXiv:2604.14430v1 Announce Type: new Abstract: We present Three-Phase Transformer (3PT), a residual-stream structural prior for decoder-only Transformers on a standard SwiGLU + RMSNorm + RoPE + GQA b

TopoDIM: One-shot Topology Generation of Diverse Interaction Modes for Multi-Agent Systems

AgentsDGX agent

arXiv:2601.10120v2 Announce Type: replace-cross Abstract: Optimizing communication topology in LLM-based multi-agent system is critical for enabling collective intelligence. Existing methods mainly re

Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms

SafetyDGX agent

arXiv:2506.09457v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs), such as Direct Preference Optimization (DPO) and Simple Preference Optimization (SimPO), have emerged as efficie

Tracking the Temporal Dynamics of News Coverage of Catastrophic and Violent Events

ResearchDGX agent

arXiv:2604.14315v1 Announce Type: new Abstract: The modern news cycle has been fundamentally reshaped by the rapid exchange of information online. As a result, media framing shifts dynamically as new

Tug-of-War within A Decade: Conflict Resolution in Vulnerability Analysis via Teacher-Guided Retrieval-Augmented Generations

ResearchDGX agent

arXiv:2604.14172v1 Announce Type: new Abstract: Large Language Models (LLMs) are essential for analyzing and addressing vulnerabilities in cybersecurity. However, among over 200,000 vulnerabilities we

Uncovering the Fragility of Trustworthy LLMs through Chinese Textual Ambiguity

Model ReleasesDGX agent

arXiv:2507.23121v2 Announce Type: replace Abstract: In this work, we study a critical research problem regarding the trustworthiness of large language models (LLMs): how LLMs behave when encountering

VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval

Model ReleasesDGX agent

arXiv:2505.20291v4 Announce Type: replace-cross Abstract: Text-to-image retrieval (T2I retrieval) remains challenging because cross-modal embeddings often behave as bags of concepts, underrepresenting

What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers

Model ReleasesDGX agent

arXiv:2604.15010v1 Announce Type: cross Abstract: When do transformers commit to a decision, and what prevents them from correcting it? We introduce extbf{prolepsis}: a transformer commits early, task

When PCOS Meets Eating Disorders: An Explainable AI Approach to Detecting the Hidden Triple Burden

Model ReleasesDGX agent

arXiv:2604.14356v1 Announce Type: new Abstract: Women with polycystic ovary syndrome (PCOS) face substantially elevated risks of body image distress, disordered eating, and metabolic challenges, yet e

Which bird does not have wings: Negative-constrained KGQA with Schema-guided Semantic Matching and Self-directed Refinement

ApplicationsDGX agent

arXiv:2604.14749v1 Announce Type: new Abstract: Large language models still struggle with faithfulness and hallucinations despite their remarkable reasoning abilities. In Knowledge Graph Question Answ

XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts

ResearchDGX agent

arXiv:2604.05242v2 Announce Type: replace Abstract: Multi-bit watermarking has emerged as a promising solution for embedding imperceptible binary messages into Large Language Model (LLM)-generated tex

XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics

Model ReleasesDGX agent

arXiv:2604.14934v1 Announce Type: new Abstract: Automatic evaluation metrics are essential for building multilingual translation systems. The common practice of evaluating these systems is averaging m

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception

SafetyDGX agent

arXiv:2510.23853v3 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used to interact with and execute tasks in dynamic environments. However, a critical yet overlook

16 Apr 2026

A closer look at how large language models trust humans: patterns and biases

ResearchDGX agent

arXiv:2504.15801v2 Announce Type: replace Abstract: As large language models (LLMs) and LLM-based agents increasingly interact with humans in decision-making contexts, understanding the trust dynamics

A Domain-Specific Language for LLM-Driven Trigger Generation in Multimodal Data Collection

Local AiDGX agent

arXiv:2604.13046v1 Announce Type: cross Abstract: Data-driven systems depend on task-relevant data, yet data collection pipelines remain passive and indiscriminate. Continuous logging of multimodal se

A Multi-Model Approach to English-Bangla Sentiment Classification of Government Mobile Banking App Reviews

SafetyDGX agent

arXiv:2604.13057v1 Announce Type: new Abstract: For millions of users in developing economies who depend on mobile banking as their primary gateway to financial services, app quality directly shapes f

A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation

Model ReleasesDGX agent

arXiv:2604.13059v1 Announce Type: new Abstract: Most dialogue-based electronic medical record (EMR) systems still behave as passive pipelines: transcribe speech, extract information, and generate the

Activation-Guided Local Editing for Jailbreaking Attacks

SafetyDGX agent

arXiv:2508.00555v2 Announce Type: replace-cross Abstract: Jailbreaking is an essential adversarial technique for red-teaming these models to uncover and patch security flaws. However, existing jailbre

Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models

ResearchDGX agent

arXiv:2604.13991v1 Announce Type: new Abstract: Large language models (LLMs) are prone to generating factually incorrect outputs. Recent work has applied conformal prediction to provide uncertainty es

Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization

TutorialsDGX agent

arXiv:2601.04442v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) have exhibited strong reasoning capabilities through chain-of-thought mechanisms that generate step-by-st

Agentic Conversational Search with Contextualized Reasoning via Reinforcement Learning

AgentsDGX agent

arXiv:2601.13115v2 Announce Type: replace Abstract: Large Language Models (LLMs) have become a popular interface for human-AI interaction, supporting information seeking and task assistance through na

AgentSPEX: An Agent SPecification and EXecution Language

AgentsDGX agent

arXiv:2604.13346v1 Announce Type: new Abstract: Language-model agent systems commonly rely on reactive prompting, in which a single instruction guides the model through an open-ended sequence of reaso

An Empirical Investigation of Practical LLM-as-a-Judge Improvement Techniques on RewardBench 2

Model ReleasesDGX agent

arXiv:2604.13717v1 Announce Type: new Abstract: LLM-as-a-judge, using a language model to score or rank candidate responses, is widely used as a scalable alternative to human evaluation in RLHF pipeli

Before the First Token: Scale-Dependent Emergence of Hallucination Signals in Autoregressive Language Models

ApplicationsDGX agent

arXiv:2604.13068v1 Announce Type: new Abstract: When do large language models decide to hallucinate? Despite serious consequences in healthcare, law, and finance, few formal answers exist. Recent work

BenGER: A Collaborative Web Platform for End-to-End Benchmarking of German Legal Tasks

Model ReleasesDGX agent

arXiv:2604.13583v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for legal reasoning requires workflows that span task design, expert annotation, model execution, and metric-bas

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size

ResearchDGX agent

arXiv:2604.13275v1 Announce Type: new Abstract: Larger language models become simultaneously better and worse at handling contextual information -- better at ignoring false claims, worse at ignoring i

Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration

SafetyDGX agent

arXiv:2604.13705v1 Announce Type: new Abstract: Fairness in language models is typically studied as a property of a single, centrally optimized model. As large language models become increasingly agen

Beyond Static Personas: Situational Personality Steering for Large Language Models

Model ReleasesDGX agent

arXiv:2604.13846v1 Announce Type: new Abstract: Personalized Large Language Models (LLMs) facilitate more natural, human-like interactions in human-centric applications. However, existing personalizat

Bi-Predictability: A Real-Time Signal for Monitoring LLM Interaction Integrity

AgentsDGX agent

arXiv:2604.13061v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in high-stakes autonomous and interactive workflows, where reliability demands continuous, multi-

Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text Detection

Model ReleasesDGX agent

arXiv:2604.13692v1 Announce Type: new Abstract: As large language models (LLMs) generate text that increasingly resembles human writing, the subtle cues that distinguish AI-generated content from huma

Bridging Compositional and Distributional Semantics: A Survey on Latent Semantic Geometry via AutoEncoder

ResearchDGX agent

arXiv:2506.20083v4 Announce Type: replace Abstract: Integrating compositional and symbolic properties into current distributional semantic spaces can enhance the interpretability, controllability, com

C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

ResearchDGX agent

arXiv:2604.13618v1 Announce Type: new Abstract: Rubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification. H

Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference

ResearchDGX agent

arXiv:2604.13634v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by letting draft tokens bypass full verification, but conventional frameworks suffer from fre

Can Large Language Models Reliably Extract Physiology Index Values from Coronary Angiography Reports?

Model ReleasesDGX agent

arXiv:2604.13077v1 Announce Type: new Abstract: Coronary angiography (CAG) reports contain clinically relevant physiological measurements, yet this information is typically in the form of unstructured

CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding

Model ReleasesDGX agent

arXiv:2604.13452v1 Announce Type: new Abstract: Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene trans

Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling

ResearchDGX agent

arXiv:2604.13054v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved rapid progress, yet their scaling behavior remains less clearly characterized and often less pred

Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs

ResearchDGX agent

arXiv:2604.13950v1 Announce Type: new Abstract: We show how causal interventions in Transformer models provide insights into English syntax by focusing on a long-standing challenge for syntactic theor

Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.13504v1 Announce Type: cross Abstract: Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging and labor-intensive process due to

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

SafetyDGX agent

arXiv:2603.27064v2 Announce Type: replace-cross Abstract: Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a ca

Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models

AgentsDGX agent

arXiv:2604.13706v1 Announce Type: new Abstract: Professional fact-checkers rely on domain knowledge and deep contextual understanding to verify claims. Large language models (LLMs) and large reasoning

CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation

Model ReleasesDGX agent

arXiv:2504.21751v4 Announce Type: replace-cross Abstract: Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components

Coherence in the brain unfolds across separable temporal regimes

Model ReleasesDGX agent

arXiv:2512.20481v4 Announce Type: replace-cross Abstract: To maintain coherence in language, the brain must satisfy key competing temporal demands: the gradual accumulation of meaning across extended

CollabCoder: Plan-Code Co-Evolution via Collaborative Decision-Making for Efficient Code Generation

Model ReleasesDGX agent

arXiv:2604.13946v1 Announce Type: cross Abstract: Automated code generation remains a persistent challenge in software engineering, as conventional multi-agent frameworks are often constrained by stat

← Previous
1…115116117118119…128
Next →