AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
24 Jul 2026

Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

AgentsDGX agent

arXiv:2607.20428v1 Announce Type: new Abstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-re

Learning to Detect UI Principle Violations via Reinforcement Learning

ApplicationsDGX agent

arXiv:2607.20690v1 Announce Type: new Abstract: Small language models and coding agents increasingly generate web front-end code, yet their outputs are typically evaluated primarily for functional cor

LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.20872v1 Announce Type: new Abstract: Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, b

MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

Model ReleasesDGX agent

arXiv:2607.21570v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or

MemTools: A Unified Research Framework for Interoperable Agent Memory

Model ReleasesDGX agent

arXiv:2607.21404v1 Announce Type: new Abstract: While memory systems are essential for agent architectures, pervasive architectural fragmentation restricts systematic research. Existing implementation

Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models

Model ReleasesDGX agent

arXiv:2607.20442v1 Announce Type: new Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in Jul

news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling

ResearchDGX agent

arXiv:2607.21284v1 Announce Type: new Abstract: Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate su

Non-Zipfian Distribution of Stopwords or Function Words and Subset Selection Models

ResearchDGX agent

arXiv:2603.04691v2 Announce Type: replace Abstract: Stopwords and function words are relatively less informative for the content of a language and more often play a structural role in a sentence. Stop

Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks

Model ReleasesDGX agent

arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single ans

Position: Natural Language Should Not Fully Replace Formal Languages

ApplicationsDGX agent

arXiv:2607.20432v1 Announce Type: new Abstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal language

PrefReward: Learning User Preference Matrix for Personalized Text Generation

TutorialsDGX agent

arXiv:2607.21067v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. H

Progressive Cramming: Reliable Token Compression and What It Reveals

ResearchDGX agent

arXiv:2607.21231v1 Announce Type: new Abstract: Token cramming compresses sequences into learned embeddings with near-perfect reconstruction, but fixed token budgets and 99% accuracy thresholds leave

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

Model ReleasesDGX agent

arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is as

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

Local AiDGX agent

arXiv:2607.20833v1 Announce Type: new Abstract: Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context wh

REGARD: Regional Affective Differences in Large Language Models

Model ReleasesDGX agent

arXiv:2607.20722v1 Announce Type: new Abstract: Large language models trained and aligned within different linguistic and regional ecosystems may frame the same political, cultural, and geopolitical e

Rushes: A Human Preference Dataset for Pluralistic Alignment

Model ReleasesDGX agent

arXiv:2607.20767v1 Announce Type: new Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collect

Sample-Efficient Learning from Agent Experience

AgentsDGX agent

arXiv:2607.21051v1 Announce Type: new Abstract: Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedbac

SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations

TutorialsDGX agent

arXiv:2607.20445v1 Announce Type: new Abstract: In conversations, human emotions are transient; however, they tend to persist across multiple utterances. For example, we rarely switch instantly betwee

Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference

ResearchDGX agent

arXiv:2607.20451v1 Announce Type: new Abstract: Semantic Field Theory (SFT) has developed from a philosophical critique of strong anti-formalist readings of language games into a proposed computationa

ShriNep@EEUCA 2026: RAKSHAK - Multi-Task DeBERTa with Rationale Distillation and Jigsaw-Augmented Training for Toxic Intent Classification

ResearchDGX agent

arXiv:2607.20450v1 Announce Type: new Abstract: This paper presents two systems for the GameTox Shared Task at the Workshop on EEUCA at ACL 2026, which requires classifying World of Tanks chat utteran

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

Model ReleasesDGX agent

arXiv:2607.09999v2 Announce Type: replace Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-catego

Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

AgentsDGX agent

arXiv:2607.20431v1 Announce Type: new Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet conv

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

Model ReleasesDGX agent

arXiv:2607.15557v4 Announce Type: replace Abstract: Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Pub

Surprisal Theory is Tautological (without Rational Grounding)

ResearchDGX agent

arXiv:2607.21574v1 Announce Type: new Abstract: Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language m

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

Model ReleasesDGX agent

arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring pro

thaulab@EEUCA 2026: Who Said What to Whom? A Targeting-Aware Neural-Symbolic Pipeline for Gaming Toxicity Detection

SafetyDGX agent

arXiv:2607.20447v1 Announce Type: new Abstract: This paper describes our system for the EEUCA 2026 Shared Task on toxicity classification in gaming chat. We implement a three-stage pipeline combining

The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning

ResearchDGX agent

arXiv:2607.20952v1 Announce Type: cross Abstract: Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assume

Token-Level Entropy Reveals Demographic Disparities in Large Language Models

SafetyDGX agent

arXiv:2501.19337v5 Announce Type: replace Abstract: A name alone measurably reshapes a language model's next-token distribution before a single token is sampled. We measure full-vocabulary Shannon ent

TopoGuard: Graph Theory Based Defenses Against Split-Knowledge Attacks on RAG

ApplicationsDGX agent

arXiv:2607.20437v1 Announce Type: new Abstract: Production Retrieval Augmented Generation (RAG) systems rely on aggregating multiple external documents to answer complex queries. However, the retrieve

Transformer-Assisted LLM-Based Source Code Summarisation: to Enable More Secure Software Development

ResearchDGX agent

arXiv:2607.20933v1 Announce Type: cross Abstract: Neural Source Code Summarisation (NSCS) aims to generate natural language summaries of source code to improve developers' and maintainers' understandi

VibeVoice-ASR-BitNet Technical Report

Local AiDGX agent

arXiv:2607.21075v1 Announce Type: cross Abstract: We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantiza

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2607.20425v1 Announce Type: new Abstract: What makes writing 'good' remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how r

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

Model ReleasesDGX agent

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

Model ReleasesDGX agent

arXiv:2607.21445v1 Announce Type: new Abstract: Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this pap

Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept

ResearchDGX agent

arXiv:2607.20995v1 Announce Type: new Abstract: Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex select

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

Model ReleasesDGX agent

arXiv:2607.21535v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models

Word meaning co-determines vowel-inherent spectral change. A corpus-based investigation of conversational Mandarin

ApplicationsDGX agent

arXiv:2607.21391v1 Announce Type: new Abstract: This study investigates vowel-inherent spectral change (VISC) in spontaneous conversational Mandarin. Using the generalized additive model and word embe

23 Jul 2026

A Multi-Dimensional Evaluation of Explainability in Media Bias Detection

SafetyDGX agent

arXiv:2607.19954v1 Announce Type: new Abstract: Detecting media bias automatically is difficult because biased framing is often subtle, yet in domains such as news analysis, accurate predictions alone

A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs

Model ReleasesDGX agent

arXiv:2607.17075v2 Announce Type: replace-cross Abstract: The advent of LLMs has significantly changed the research on privacy policy and data compliance analysis by enabling tasks that previously req

Abstraction Induces the Brain Alignment of Language and Speech Models

SafetyDGX agent

arXiv:2602.04081v2 Announce Type: replace Abstract: Research has repeatedly demonstrated that intermediate hidden states extracted from large language models and speech audio models predict measured b

AugAbEx: Bridging Abstractive and Extractive Legal Summarization

ApplicationsDGX agent

arXiv:2511.12290v2 Announce Type: replace Abstract: Automatic summarization of legal judgments liberates law professionals from heavy cognitive burden due to the complexity of the language, context-se

Back to Back with a Copy: A Computational Analysis of AI-Generated Visual Contemporary Art Pastiches

SafetyDGX agent

arXiv:2607.20127v1 Announce Type: new Abstract: The aim of this paper is twofold. First, it investigates whether newer generative models are getting better at pastiching contemporary artworks. Second,

Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

Model ReleasesDGX agent

arXiv:2607.19747v1 Announce Type: new Abstract: As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream gen

Building Fast, Evaluating Slow: Pipeline Choices Dominate Autointerpretability Score Variance

ResearchDGX agent

arXiv:2607.19386v1 Announce Type: cross Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores. In this evaluation pipeline, a langua

D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

Model ReleasesDGX agent

arXiv:2607.19834v1 Announce Type: new Abstract: With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

ResearchDGX agent

arXiv:2607.19932v1 Announce Type: new Abstract: Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

ResearchDGX agent

arXiv:2607.19848v1 Announce Type: new Abstract: There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity rem

Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation: A Comparative Study

Model ReleasesDGX agent

arXiv:2502.16022v3 Announce Type: replace Abstract: OpenNotes gives patients access to their EHR notes, but dense medical jargon limits comprehension. We evaluate closed-source and open-source LLMs fo

Exposure is Optional: Learning Unlike Coordination in Language Models

ResearchDGX agent

arXiv:2607.20251v1 Announce Type: new Abstract: Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics.

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis

Model ReleasesDGX agent

arXiv:2607.09530v2 Announce Type: replace Abstract: We introduce Freya-TTS, a compact, tokenizer-free, Turkish-first text-to-speech model designed for highly reliable and efficient conversational synt

Gotta Catch them all: the modes of Sycophancy

ResearchDGX agent

arXiv:2607.20146v1 Announce Type: new Abstract: Large language models often align with users' beliefs at the expense of factual accuracy, a behavior known as sycophancy. Prior mechanistic studies larg

H^2SD: Hybrid Hindsight Self-Distillation

ResearchDGX agent

arXiv:2607.18955v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar traject

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Model ReleasesDGX agent

arXiv:2607.20219v1 Announce Type: new Abstract: Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing

It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

Model ReleasesDGX agent

arXiv:2607.18232v2 Announce Type: replace Abstract: Users frequently express their beliefs to large language models (LLMs). In some situations, the LLM should accept these contextual beliefs as true.

LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization

Local AiDGX agent

arXiv:2407.00740v2 Announce Type: replace Abstract: As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, suc

Lightweight Person-Place Relation Extraction from Historical Newspapers with Dependency Graphs and Proximity Features

TutorialsDGX agent

arXiv:2607.19718v1 Announce Type: new Abstract: The HIPE-2026 shared task introduces person-place relation extraction from multilingual historical newspapers as a new evaluation track, classifying the

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Model ReleasesDGX agent

arXiv:2607.20410v1 Announce Type: new Abstract: Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local va

Meta-Learning Preferences for Multilingual LLM Alignment

SafetyDGX agent

arXiv:2607.13315v2 Announce Type: replace Abstract: Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual sett

Multi-Mask Diffusion Language Models for Few-Step Generation

ResearchDGX agent

arXiv:2607.19686v1 Announce Type: new Abstract: Masked diffusion models (MDMs) are a promising family of language generators, but achieving high-quality few-step generation remains challenging. In MDM

Notes to Self: Can LLMs Benefit from Experiential Abstractions?

ResearchDGX agent

arXiv:2607.20372v1 Announce Type: new Abstract: Humans distill experience into reusable abstractions, e.g., strategies and cautionary reminders, and apply them to gradually solve problems more effecti

← Previous
1…2021222324…129
Next →