AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
3 Aug 2026

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges

SafetyDGX agent

arXiv:2607.28636v1 Announce Type: new Abstract: LLMs increasingly serve as automated judges, but their judgments remain vulnerable to cognitive biases. Existing mitigations mostly rely on prompt-drive

Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation

Model ReleasesDGX agent

arXiv:2607.29250v1 Announce Type: new Abstract: Small language models (SLMs) are attractive for agentic deployment due to low latency, reduced cost, and on-device privacy, yet they struggle with tool-

Demystifying Entropy-based Selection for Chain-of-Thought Compression in Large Reasoning Models

ResearchDGX agent

arXiv:2607.28707v1 Announce Type: new Abstract: Entropy-based pruning has been proposed as an effective method for compressing Chain-of-Thought (CoT) reasoning with negligible accuracy loss. We test t


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives

Model ReleasesDGX agent

arXiv:2607.29188v1 Announce Type: new Abstract: Migrants traversing geographically distinct routes such as the Trans-Saharan and Balkan corridors often recount strikingly parallel lived experiences: p

Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

ResearchDGX agent

arXiv:2603.24917v3 Announce Type: replace Abstract: Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across seque

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

ResearchDGX agent

arXiv:2607.29484v1 Announce Type: new Abstract: Interventional data is widely regarded as the gold standard for teaching models causal reasoning. We test this assumption in a fully controlled syntheti

Evolving language compositionality in a frequency-structured meaning space

ResearchDGX agent

arXiv:2607.29642v1 Announce Type: new Abstract: The iterated learning model was introduced to investigate language evolution: the way in which the characteristic properties of human languages have bee

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models

SafetyDGX agent

arXiv:2607.29079v1 Announce Type: new Abstract: Training-free acceleration makes diffusion-based multimodal large language models (dMLLMs) more deployable, but it may silently change generated content

FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale

ResearchDGX agent

arXiv:2601.22146v3 Announce Type: replace Abstract: Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised 'predict the next word' object

From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts

ApplicationsDGX agent

arXiv:2607.29044v1 Announce Type: new Abstract: Inline notes and collected commentaries are important forms of scholarly communication that evolved within the Confucian exegetical tradition, yet have

GoldenRetriever: Non-Interactive Homomorphic Encrypted Retrieval for Privacy-Preserving RAG

ResearchDGX agent

arXiv:2607.29019v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances large language models by incorporating external knowledge, but existing pipelines typically operate on p

HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution

AgentsDGX agent

arXiv:2607.13683v2 Announce Type: replace Abstract: Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is go

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

Model ReleasesDGX agent

arXiv:2607.29196v1 Announce Type: new Abstract: Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, track

Imbalanced Data Clustering via Targeted Data Augmentation Using GMM and LLM

ResearchDGX agent

arXiv:2607.28635v1 Announce Type: new Abstract: In Natural Language Processing (NLP), dealing with underrepresented topics is challenging, especially in unsupervised tasks where clustering might not a

Know It, Act on It: Investigating Memory Utilization in LLM Personalization

AgentsDGX agent

arXiv:2607.29433v1 Announce Type: new Abstract: As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utiliz

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

ResearchDGX agent

arXiv:2607.29211v1 Announce Type: new Abstract: Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-soun

Language Models Agree With Each Other, Not With Readers

Model ReleasesDGX agent

arXiv:2607.29274v1 Announce Type: cross Abstract: Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact o

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

SafetyDGX agent

arXiv:2607.29185v1 Announce Type: new Abstract: Reward models (RMs) are central to aligning large language models with human preferences via reinforcement learning. Although traditional scalar RMs ena

Learning Stateful Predictive Knowledge From Experience

SafetyDGX agent

arXiv:2607.28638v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

Model ReleasesDGX agent

arXiv:2607.29125v1 Announce Type: new Abstract: Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and

Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents

AgentsDGX agent

arXiv:2607.28651v1 Announce Type: cross Abstract: Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an ext

Mixture-of-Translators: Translating KV Caches Across Heterogeneous Large Language Models

AgentsDGX agent

arXiv:2607.28979v1 Announce Type: new Abstract: Heterogeneous Large Language Model (LLM) systems increasingly rely on shared contexts, retrieved evidence, and multi-agent dialogue histories, yet their

PTP: Previous-Token Prediction based LLM Inversion for Near-Exact Prompt Reconstruction

ResearchDGX agent

arXiv:2607.29378v1 Announce Type: new Abstract: Large language models (LLMs) generate text by auto-regressively sampling the next token. This inherently leads to a many-to-many mapping between prompts

ResKV: Reconstructing Omitted Attention Contributions for Fixed-Budget KV Cache Compression

ResearchDGX agent

arXiv:2607.29591v1 Announce Type: new Abstract: KV cache compression is essential for efficient long-context inference. Existing eviction methods permanently discard unselected tokens and consequently

Self-Supervised Skill Optimization

AgentsDGX agent

arXiv:2607.28777v1 Announce Type: new Abstract: Agent skills provide frozen large language model (LLM) agents with reusable procedural guidance, and recent work shows that such skills can be optimized

Studying quantization trade-offs for efficient inference deployment in machine translation

HardwareDGX agent

arXiv:2607.29397v1 Announce Type: new Abstract: Deploying large language models in realistic server environments poses challenges, as the system needs to provide high-quality responses with low latenc

Sycophancy Undermines Epistemic Vigilance in Cooperative Vision-Language Tasks

ResearchDGX agent

arXiv:2607.29585v1 Announce Type: new Abstract: To maintain common ground in cooperative conversation, humans iteratively update their beliefs as conversation participants share new information; parti

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

SafetyDGX agent

arXiv:2607.28680v1 Announce Type: new Abstract: Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on

The Checking Problem: What must be true before AI ships in a regulated firm

ApplicationsDGX agent

arXiv:2607.28666v1 Announce Type: new Abstract: Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of

The Morphological Core of Dungan: A Two-Dialect Finite-State Model and a Multi-Genre Evaluation

Model ReleasesDGX agent

arXiv:2607.28766v1 Announce Type: new Abstract: Dungan, a Sinitic language of Central Asia written in a Cyrillic-based script, is described in detail in the grammatical literature, yet the quantitativ

To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions

TutorialsDGX agent

arXiv:2607.28643v1 Announce Type: cross Abstract: Automating facilitation in online discussions is a long-standing social concern given the increasing time we spend on online spaces and the failure of

Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

ResearchDGX agent

arXiv:2607.28906v1 Announce Type: new Abstract: Sycophancy refers to the tendency for large language models (LLMs) to match user beliefs at the cost of factual correctness, thereby undermining model r

Tokenizer-Agnostic Engram Module

Model ReleasesDGX agent

arXiv:2607.29065v1 Announce Type: new Abstract: Deepseek's Engram, a conditional memory module, was introduced to trade-off storage versus reasoning in large language models. However, the module relie

Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR

ResearchDGX agent

arXiv:2607.09598v2 Announce Type: replace Abstract: Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologica

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

ResearchDGX agent

arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observ

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

HardwareDGX agent

arXiv:2607.29678v1 Announce Type: new Abstract: LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, w

TransMem: Transforming Hidden States into Memory for Large Language Models

TutorialsDGX agent

arXiv:2607.29032v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly operate over long interaction histories, where effective reasoning requires identifying and exploiting

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

SafetyDGX agent

arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods,

Zero-Mem: Zero-Token Memory Operations for LLM Agents

Local AiDGX agent

arXiv:2607.29377v1 Announce Type: new Abstract: LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermed

ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification

ResearchDGX agent

arXiv:2607.28637v1 Announce Type: new Abstract: This paper presents our system for the CHiPSAL 2026 shared task on multimodal hate speech and sentiment detection in Nepali memes. We address both subta

31 Jul 2026

A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding

ResearchDGX agent

arXiv:2607.27735v1 Announce Type: new Abstract: Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafti

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports

Model ReleasesDGX agent

arXiv:2601.15297v3 Announce Type: replace Abstract: Reliable question answering over long institutional documents requires more than topical retrieval: a system must localize the exact passage that su

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

Model ReleasesDGX agent

arXiv:2607.27393v1 Announce Type: new Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cul

AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026

SafetyDGX agent

arXiv:2607.27228v1 Announce Type: new Abstract: Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences ar

AI systems and the reproduction of (standard) language ideologies in World Englishes

ApplicationsDGX agent

arXiv:2607.28528v1 Announce Type: new Abstract: The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what cou

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

ResearchDGX agent

arXiv:2607.28617v1 Announce Type: cross Abstract: System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout com

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Model ReleasesDGX agent

arXiv:2607.28618v1 Announce Type: new Abstract: Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems pr

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

Model ReleasesDGX agent

arXiv:2607.17219v2 Announce Type: replace Abstract: Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free predict

AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

Model ReleasesDGX agent

arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability

AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure

ResearchDGX agent

arXiv:2607.27611v1 Announce Type: new Abstract: Corporate annual reports contain weakly structured evidence about foreign-exchange risk management, derivative use, natural hedging, and explicit non-us

Baikal: Structured Search for Deep Research over Data Lakes

Model ReleasesDGX agent

arXiv:2607.27726v1 Announce Type: cross Abstract: Deep research over data lakes requires an LLM agent to investigate evidence across thousands of heterogeneous tables and passages to synthesize a repo

Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models

AgentsDGX agent

arXiv:2607.27512v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among int

Benchmarking LLM Competence on Logical Inference over Probability Operators

Model ReleasesDGX agent

arXiv:2607.27405v1 Announce Type: new Abstract: Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty

Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation

ResearchDGX agent

arXiv:2607.28439v1 Announce Type: new Abstract: Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluatin

Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation

Model ReleasesDGX agent

arXiv:2607.27816v1 Announce Type: new Abstract: Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversation

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

SafetyDGX agent

arXiv:2607.27851v1 Announce Type: new Abstract: Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience

Beyond Sentiment: Structured Information Extraction from Financial News

Model ReleasesDGX agent

arXiv:2607.28496v1 Announce Type: new Abstract: Financial sentiment analysis has become a standard component in news-driven stock prediction, yet it reduces rich, multi-dimensional news articles to a

Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories

Model ReleasesDGX agent

arXiv:2607.27595v1 Announce Type: new Abstract: Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-pa

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

SafetyDGX agent

arXiv:2607.27366v1 Announce Type: new Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanit

CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance

ApplicationsDGX agent

arXiv:2607.28292v1 Announce Type: new Abstract: Large Language Models (LLMs) deployed in dynamic financial environments face a critical challenge: maintaining factual accuracy as market conditions, re

← Previous
1…1112131415…128
Next →