AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
7 Aug 2026

Building Open-Retrieval Conversational Question Answering Systems by Generating Synthetic Data and Decontextualizing User Questions

ApplicationsDGX agent

arXiv:2507.04884v2 Announce Type: replace Abstract: We consider open-retrieval conversational question answering (OR-CONVQA), an extension of question answering where system responses need to be (i) a

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

AgentsDGX agent

arXiv:2608.06352v1 Announce Type: cross Abstract: Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable

Causal Episodic Memory for Feedback-Driven Agent Repair

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.05906v1 Announce Type: new Abstract: LLM agents that repair failures often discard successful corrections, forcing later episodes to rediscover similar solutions. We study whether finalized

Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies

SafetyDGX agent

arXiv:2608.05993v1 Announce Type: new Abstract: Much clinical value is conveyed not through structured records but through communication: exchanges in which patients describe symptoms, clinicians reas

CNM-BERT: A Drop-In Structural Embedding for Chinese Characters via Ideographic Description Sequences

Model ReleasesDGX agent

arXiv:2608.05167v1 Announce Type: new Abstract: Token-based encoders like BERT treat Chinese characters as atomic identifiers, ignoring their recursive orthographic structure. Consequently, models rel

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

Model ReleasesDGX agent

arXiv:2608.05166v1 Announce Type: new Abstract: We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our wo

Constraint-First Reasoning: A Training-Free Protocol for Exploiting Answer-Space Constraints in Mathematical Problem Solving

ResearchDGX agent

arXiv:2608.05254v1 Announce Type: new Abstract: Large language models can derive a plausible mathematical object yet still violate explicit requirements--for example, by omitting a modular reduction,

ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control

Model ReleasesDGX agent

arXiv:2608.05169v1 Announce Type: new Abstract: Long-form story generation requires models to preserve narrative consistency across extended contexts, yet existing prompting-based methods often accumu

CPC-CMS: Cognitive Pairwise Comparison Classification Model Selection Framework for Document-level Sentiment Analysis

ResearchDGX agent

arXiv:2507.14022v2 Announce Type: replace Abstract: This study proposes the Cognitive Pairwise Comparison Classification Model Selection (CPC-CMS) framework for document-level sentiment analysis. The

Cross-Architecture Steering Transfer in Language Models: A Systematic Empirical Study

Model ReleasesDGX agent

arXiv:2608.05164v1 Announce Type: new Abstract: Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but wh

DBLAST: Dependent Block Drafting for Stochastic Speculative Decoding

ResearchDGX agent

arXiv:2608.05448v1 Announce Type: new Abstract: Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

SafetyDGX agent

arXiv:2608.06141v1 Announce Type: new Abstract: This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and educati

Decomposed Entailment for Factuality Checking and Hallucination Detection

ResearchDGX agent

arXiv:2608.05823v1 Announce Type: new Abstract: The reliability of Large Language Models (LLMs) is often compromised by factual inconsistencies, including hallucinations---cases where generated conten

Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness

SafetyDGX agent

arXiv:2608.05510v1 Announce Type: new Abstract: Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promisi

EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding

Local AiDGX agent

arXiv:2608.05303v1 Announce Type: cross Abstract: On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. A primary bottleneck is external memory

Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding

Model ReleasesDGX agent

arXiv:2608.05832v1 Announce Type: new Abstract: Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

Model ReleasesDGX agent

arXiv:2608.06022v1 Announce Type: new Abstract: Epitopes determine where antibodies bind antigens and shape downstream therapeutic properties such as functional blockade and escape resistance, making

Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation

Model ReleasesDGX agent

arXiv:2608.05353v1 Announce Type: new Abstract: LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate recor

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

SafetyDGX agent

arXiv:2608.05446v1 Announce Type: cross Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, invoke tools, verify outcomes, and reuse ex

Example-Guided Prompting for Document-Level Text Simplification

ResearchDGX agent

arXiv:2608.05447v1 Announce Type: new Abstract: Document-level text simplification requires large language models (LLMs) to rewrite complex documents while preserving meaning, readability, and discour

FOCUS: Decoupling Expert Personas in LLMs to Enhance Domain Expert Capabilities

ApplicationsDGX agent

arXiv:2608.05611v1 Announce Type: new Abstract: Large Language Models (LLMs) can exhibit diverse personas, and activating expert personas has been shown to improve domain expertise and task accuracy.

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs

Model ReleasesDGX agent

arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, l

GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers

SafetyDGX agent

arXiv:2608.05478v1 Announce Type: cross Abstract: Graphical Abstracts (GAs) visually summarize the key findings of academic papers, playing a crucial role in facilitating the understanding of research

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

TutorialsDGX agent

arXiv:2608.05759v1 Announce Type: new Abstract: Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challe

Human-Like Anaphor Resolution in Large Language Models

Model ReleasesDGX agent

arXiv:2608.05630v1 Announce Type: new Abstract: Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science

Integrating Human Linguistic Insights into AI: Theory-Driven Representation for Multilingual Text-to-Speech

ResearchDGX agent

arXiv:2204.07228v2 Announce Type: replace Abstract: This paper explores the integration of human linguistic insights into multilingual text-to-speech (TTS) systems by evaluating the Featurally Undersp

LangChoiceBench: Measuring and Explaining Programming-Language Choice in LLMs

Model ReleasesDGX agent

arXiv:2608.06041v1 Announce Type: cross Abstract: Large language models (LLMs) have been shown to exhibit strong Python preferences when generating project-level code, but there is currently no system

LELA: an LLM-based Entity Linking Approach with Zero-Shot Domain Adaptation

ResearchDGX agent

arXiv:2601.05192v2 Announce Type: replace Abstract: Entity linking (mapping ambiguous mentions in text to entities in a knowledge base) is a foundational step in tasks such as knowledge graph construc

M^3R-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding

Model ReleasesDGX agent

arXiv:2608.05817v1 Announce Type: new Abstract: Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visu

Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs

SafetyDGX agent

arXiv:2510.03997v2 Announce Type: replace Abstract: Understanding how patients perceive their physicians is essential to improving trust, communication, and satisfaction. Patients increasingly consult

Mapping Similarity Spaces across Embedding Models with Synthetic Query Probing

TutorialsDGX agent

arXiv:2608.05857v1 Announce Type: new Abstract: Retrieval-Augmented Generation systems rely on similarity scores to retrieve relevant content, yet scores are not directly comparable across embedding m

Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation

SafetyDGX agent

arXiv:2608.05726v1 Announce Type: new Abstract: Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluati

MoCA: Implicit Social Context Analysis

Model ReleasesDGX agent

arXiv:2608.05825v1 Announce Type: new Abstract: Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through ind

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

SafetyDGX agent

arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep est

NeSy-RAG: Neuro-Symbolic RAG for Explainable Question Answering

Model ReleasesDGX agent

arXiv:2608.06292v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. H

On-Policy Delta Distillation for Multilingual Math Reasoning

SafetyDGX agent

arXiv:2608.05802v1 Announce Type: new Abstract: On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingu

Persona-Pruner: Sculpting Lightweight Models for Role-Playing

ApplicationsDGX agent

arXiv:2606.14695v2 Announce Type: replace-cross Abstract: Language Models (LMs) have shown remarkable potential as role-playing chatbots, delivering consistent, stylized interactions when given a spec

PolyAlign: Conditional Human-Distribution Alignment

SafetyDGX agent

arXiv:2606.13227v2 Announce Type: replace Abstract: Post-training methods such as supervised fine-tuning (SFT) and preference optimization typically align language models toward a single global assist

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

Model ReleasesDGX agent

arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden

Predicting Social Media User Actions: A Hybrid Approach for Common and Rare Behavior Prediction on Bluesky

ResearchDGX agent

arXiv:2511.17241v2 Announce Type: replace Abstract: Understanding and predicting user behavior on social media platforms is crucial for content recommendation and platform design. While existing appro

Predicting Task Difficulty Without Rollouts

AgentsDGX agent

arXiv:2608.05797v1 Announce Type: cross Abstract: Task difficulty dictates an agent's likelihood of success, and estimating it without rollouts means forecasting this directly from a task description

QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding

ResearchDGX agent

arXiv:2608.05326v1 Announce Type: cross Abstract: Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work

Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

ResearchDGX agent

arXiv:2608.05660v1 Announce Type: cross Abstract: As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning ha

RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates

ResearchDGX agent

arXiv:2608.05154v1 Announce Type: new Abstract: Rotary positional encoding (RoPE) is a core component of modern language models and has been extended to multimodal LLMs through multidimensional varian

Robust Native Language Identification through Agentic Decomposition

Model ReleasesDGX agent

arXiv:2509.16666v2 Announce Type: replace Abstract: Large language models (LLMs) often achieve high performance in native language identification (NLI) benchmarks by leveraging superficial contextual

Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents

AgentsDGX agent

arXiv:2608.06171v1 Announce Type: new Abstract: Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks. We measure six observation modes across

RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer

SafetyDGX agent

arXiv:2608.06347v1 Announce Type: new Abstract: Multilingual reasoning transfer is crucial for extending reasoning capabilities of large language models (LLMs) beyond high-resource languages. On-polic

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

ResearchDGX agent

arXiv:2608.06310v1 Announce Type: cross Abstract: Recent advances in reward modeling show a paradigm shift from discriminative reward models to generative reward models. However, despite their strong

Safe Evolution with Circuit Anchors

SafetyDGX agent

arXiv:2608.05158v1 Announce Type: new Abstract: In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential fun

Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs

Model ReleasesDGX agent

arXiv:2608.05156v1 Announce Type: new Abstract: Post-training of large language models optimizes only parameters, while inference-time procedural scaffolds are typically designed independently of para

SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters

Model ReleasesDGX agent

arXiv:2608.05161v1 Announce Type: new Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining rema

Shrinking the Generation-Verification Gap with Weak Verifiers

Model ReleasesDGX agent

arXiv:2506.18203v3 Announce Type: replace Abstract: Verifiers can improve language model capabilities by scoring and ranking responses from generated candidates. Currently, high-quality verifiers are

Sparse Mutual Information Graph Averaging for Improving Random Indexing Embeddings

ResearchDGX agent

arXiv:2608.05724v1 Announce Type: new Abstract: Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on

STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts

TutorialsDGX agent

arXiv:2602.14265v3 Announce Type: replace Abstract: Inference-Time-Compute (ITC) methods like Best-of-n and Tree-of-Thoughts are meant to produce output candidates that are both high-quality and diver

Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges

SafetyDGX agent

arXiv:2405.15604v4 Announce Type: replace Abstract: Text generation has become more accessible than ever, and the growing interest in these systems, especially those using large language models, has s

The Bitter Lesson of Tool Calling

Model ReleasesDGX agent

arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by

The interface of intonation and lexical tone: Boundary phenomena in Mandarin varieties

ResearchDGX agent

arXiv:2608.05364v1 Announce Type: new Abstract: This chapter explores the intricate interplay between intonation and tone in Mandarin Chinese varieties, focusing on f0, the primary acoustic cue for bo

The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents

AgentsDGX agent

arXiv:2608.05884v1 Announce Type: cross Abstract: Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks.

Training-Free Token-Level Steering for LLM Personalized Co-Writing

ResearchDGX agent

arXiv:2608.06069v1 Announce Type: new Abstract: While Large Language Models (LLMs) show great promise for personalization, they often lack specialized domain knowledge. Conventional solutions like fin

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

ResearchDGX agent

arXiv:2608.05576v1 Announce Type: new Abstract: When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizab

← Previous
1…34567…128
Next →