AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
7 Jul 2026

Why teaching resists automation in an AI-inundated era: Human judgment, non-modular work, and the limits of delegation

ApplicationsDGX agent

arXiv:2604.07285v2 Announce Type: replace Abstract: Debates about artificial intelligence (AI) in education often portray teaching as a modular and procedural job that can increasingly be automated or

WPG-MoE: Weak-Prior-Guided Dense Mixture-of-Experts for User-Level Social Media Depression Detection

Local AiDGX agent

arXiv:2607.04350v1 Announce Type: new Abstract: Online social media posts provide scalable signals for early depression screening, and recent studies mainly improve pre-classification evidence through

Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.04640v1 Announce Type: new Abstract: We study how correctness is assembled inside aligned language models, not only whether the final answer is right. Using layer-wise difference-in-differe

You Frame It: How Conceptual Representations Shape LLM Detection and Reasoning about Antisemitism

ResearchDGX agent

arXiv:2607.04945v1 Announce Type: new Abstract: LLMs enable the integration of external conceptual resources at inference time, creating new opportunities for detecting ideologically and historically

3 Jul 2026

AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG

Model ReleasesDGX agent

arXiv:2602.19127v2 Announce Type: replace Abstract: With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop rea

AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs

ResearchDGX agent

arXiv:2601.22710v2 Announce Type: replace-cross Abstract: Modern LLMs are increasingly accessed via black-box APIs, requiring users to transmit sensitive prompts, outputs, and fine-tuning data to exte

AthDGC: An Open Diachronic Greek Treebank with Indo-European Parallels

SafetyDGX agent

arXiv:2606.15510v2 Announce Type: replace Abstract: AthDGC ('Athens-PROIEL') is an open, end-to-end workflow and dataset. It is, to the best of our knowledge, the first openly licensed dependency-pars

Audio-Based Understanding of Audiobook Narration Appeal

ResearchDGX agent

arXiv:2607.02473v1 Announce Type: new Abstract: Narration is central to the audiobook listening experience, shaping how listeners engage with and understand the content. This work explores how narrati

BamiBERT: A New BERT-based Language Model for Vietnamese

ResearchDGX agent

arXiv:2607.02259v1 Announce Type: new Abstract: In this paper, we introduce BamiBERT, a new BERT-based pre-trained language model for Vietnamese that addresses key limitations of PhoBERT -- the curren

Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation

Model ReleasesDGX agent

arXiv:2607.02182v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit remarkable reasoning capabilities, but their task-specific fine-tuning is notoriously plagued by overconfidence,

Beyond Pixel Diffs: Benchmarking Image Change Captioning for Web UI Visual Regression Testing

Model ReleasesDGX agent

arXiv:2607.01728v1 Announce Type: cross Abstract: Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interfa

Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework

Model ReleasesDGX agent

arXiv:2607.01581v1 Announce Type: new Abstract: The capacity of Large Language Models (LLMs) to reason about pedagogical intent within instructional communication remains underexplored, particularly i

Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsing

AgentsDGX agent

arXiv:2607.01964v1 Announce Type: new Abstract: Rewriting inputs to improve frozen downstream models has become a common strategy in modern NLP pipelines. Prior work on incremental dialogue discourse

BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2607.01600v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

ResearchDGX agent

arXiv:2607.01538v1 Announce Type: new Abstract: Language models (LMs) raise an intriguing alternative to vector-based retrieval: conditioning on an in-context corpus and directly generating a relevant

CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning

Local AiDGX agent

arXiv:2607.02262v1 Announce Type: new Abstract: Reasoning Language Models (RLMs) have significantly improved performance on complex tasks by extending the reasoning chain. However, these chains are pr

Comparing Architectures for Supervised Political Scaling

ResearchDGX agent

arXiv:2607.01464v1 Announce Type: new Abstract: Text scaling, the task of positioning political actors on an ideological scale, is a fundamental task in political analysis. To ease the need for manual

Denser neq Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Model ReleasesDGX agent

arXiv:2607.01763v1 Announce Type: cross Abstract: Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy

Do LLMs Truly Generalize in the Molecular Domain? A Perturbation-Based Analysis

ResearchDGX agent

arXiv:2607.01800v1 Announce Type: cross Abstract: Large Language Models (LLMs) have recently shown promise in molecular discovery, yet a gap remains between their probabilistic nature over discrete se

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models

Model ReleasesDGX agent

arXiv:2607.02007v1 Announce Type: new Abstract: Large language models now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single dis

FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning

AgentsDGX agent

arXiv:2607.01440v1 Announce Type: new Abstract: Faithful reasoning is essential in medicine, where clinical decisions require transparent justification grounded in reliable evidence. Current medical L

From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages

Model ReleasesDGX agent

arXiv:2607.01502v1 Announce Type: new Abstract: Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space mo

Gender Differences in Research Topic and Method Selection in Library and Information Science: Perspectives from Three Top Journals

ResearchDGX agent

arXiv:2607.01828v1 Announce Type: cross Abstract: Research in the social sciences has shown that there are gender differences in the selection of research methods, with women often opting for qualitat

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

Model ReleasesDGX agent

arXiv:2607.02079v1 Announce Type: new Abstract: We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art perfor

HNSW with Accuracy Guarantees Using Graph Spanners -- A Technical Report

Model ReleasesDGX agent

arXiv:2607.02338v1 Announce Type: cross Abstract: Hierarchical Navigable Small World (HNSW) graphs serve as the industry standard due to their logarithmic complexity and strong empirical performance.

HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generation

Model ReleasesDGX agent

arXiv:2607.02381v1 Announce Type: new Abstract: This paper describes the participation of HULAT2-UC3M in the Spanish track of MER-TRANS 2026, a shared task on multilingual Easy-to-Read translation. Th

Know Your Source: A Public Knowledge Store for Media Background Checks

ApplicationsDGX agent

arXiv:2607.02383v1 Announce Type: new Abstract: LLM-based retrieval-augmented generation (RAG) is increasingly used for automated fact-checking (AFC) and related tasks. By grounding LLM outputs in ret

Language Models as Measurement Apparatus for Culture

AgentsDGX agent

arXiv:2607.02459v1 Announce Type: new Abstract: Language models are increasingly used to quantify cultural phenomena, but what makes such measurement distinctively cultural? This paper argues that NLP

Large language models reshape the language of science

ApplicationsDGX agent

arXiv:2504.12317v2 Announce Type: replace Abstract: Scientific language is a central infrastructure of knowledge production, but it remains unclear whether large language models (LLMs) are altering no

LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models

Model ReleasesDGX agent

arXiv:2504.02327v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive acce

Multi-Objective Exploration and Preference Optimization via Mutual Information

SafetyDGX agent

arXiv:2607.01392v1 Announce Type: new Abstract: Aligning large language models with diverse and heterogeneous human values requires multi-objective alignment methods to effectively trade off conflicti

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

ResearchDGX agent

arXiv:2607.01960v1 Announce Type: new Abstract: In this paper, we describe NAVER LABS Europe's submission to the instruction-following speech processing short track at IWSLT 2026. We participate again

Non-synchronism in Global Usage of Research Methods in Library and Information Science from 1990 to 2019

TutorialsDGX agent

arXiv:2607.01833v1 Announce Type: cross Abstract: The global development of Library and Information Science (LIS) is influenced by various factors such as the economy, society, culture, discipline, tr

On the Limits of Steering Vectors for Preference-Aligned Generation

Model ReleasesDGX agent

arXiv:2607.01802v1 Announce Type: new Abstract: Steering vectors have emerged as a promising approach to controlled text generation, offering interpretable, training-free mechanisms for shaping model

On the Role of Directionality in Structural Generalization

ResearchDGX agent

arXiv:2607.02307v1 Announce Type: new Abstract: Several SLOG test categories explicitly involve directional distinctions (modifier position shifts, argument extraction positions), yet AM-Parser, the p

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation

Model ReleasesDGX agent

arXiv:2607.01883v1 Announce Type: new Abstract: Code is the medium through which large language models generate structured artifacts: charts, scientific figures, vector graphics, CAD models, 3D scenes

Parameter Golf: What Really Works?

Model ReleasesDGX agent

arXiv:2607.01517v1 Announce Type: new Abstract: How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which particip

PARTREP: Learning What to Repeat for Decoder-only LLMs

ResearchDGX agent

arXiv:2607.01792v1 Announce Type: new Abstract: While decoder-only LLMs excel at a vast array of natural language tasks, it suffers from an asymmetric information flow induced by causal attention: lat

Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

Model ReleasesDGX agent

arXiv:2506.12311v4 Announce Type: replace Abstract: Text-to-speech (TTS) for Modern Hebrew is challenged by the language's orthographic complexity, with existing solutions ignoring underspecified phon

Probing Spectrum-Like Organization of States of Mind in Transformer Representation Spaces

ResearchDGX agent

arXiv:2512.22227v3 Announce Type: replace Abstract: We investigate whether graded states of mind form spectrum-like structure in transformer representation spaces. To do so, we construct a dataset of

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators

ResearchDGX agent

arXiv:2607.01602v1 Announce Type: new Abstract: SRAM-based FPGAs provide an attractive platform for energy- and latency-constrained CNN inference at the network edge, yet transient faults can lead to

Recursive Models for Long-Horizon Reasoning

AgentsDGX agent

arXiv:2603.02112v2 Announce Type: replace-cross Abstract: Modern language models reason within bounded context, an inherent constraint that poses a fundamental barrier to long-horizon reasoning. We id

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

ResearchDGX agent

arXiv:2607.01733v1 Announce Type: new Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recogniti

RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules

ResearchDGX agent

arXiv:2607.01293v1 Announce Type: new Abstract: We present RuleChef, a framework that uses large language models (LLMs) to generate executable rules for NLP tasks such as text classification, Named En

RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

Model ReleasesDGX agent

arXiv:2607.01388v1 Announce Type: new Abstract: Multi-step symbolic reasoning is essential for robust financial analysis, yet most benchmarks neglect intermediate reasoning steps. FINCHAIN introduced

Scaling Latent Reasoning via Looped Language Models

ResearchDGX agent

arXiv:2510.25741v5 Announce Type: replace Abstract: Modern LLMs are trained to 'think' primarily via explicit text generation, such as chain-of-thought (CoT), which defers reasoning to post-training a

Self-Supervised Test-Time Tuning for Packet Loss Concealment

ResearchDGX agent

arXiv:2607.01823v1 Announce Type: cross Abstract: Packet loss concealment (PLC) reconstructs audio packets that are missing at the receiver, usually with a trained model whose parameters remain fixed

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics

Model ReleasesDGX agent

arXiv:2510.09517v2 Announce Type: replace Abstract: Despite rapid advances in large language models (LLMs), statistical reasoning remains underrepresented in existing LLM benchmarks, which often do no

The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing

ResearchDGX agent

arXiv:2607.02416v1 Announce Type: new Abstract: Natural Language Processing (NLP) has traditionally been published in its core disciplinary venues like ACL. However, advances in Large Language Models

The Grammar Does the Work: Functional vs. Lexical Dependency Length Minimization Across Universal Dependencies

ResearchDGX agent

arXiv:2607.01899v1 Announce Type: new Abstract: Dependency length minimization (DLM) is a well-documented processing universal, but previous studies report a single mean dependency distance (MDD) per

Towards a Phonology-Informed Evaluation of Multilingual TTS

Model ReleasesDGX agent

arXiv:2607.01965v1 Announce Type: new Abstract: Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words fro

Towards Robustness against Typographic Attack with Training-free Concept Localization

Model ReleasesDGX agent

arXiv:2607.02494v1 Announce Type: cross Abstract: Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Model

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

ResearchDGX agent

arXiv:2607.02214v1 Announce Type: new Abstract: Instruction tuning for speech language models (SLMs) is substantially more challenging than for text-based large language models (LLMs), as it requires

Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words

ResearchDGX agent

arXiv:2607.02002v1 Announce Type: new Abstract: Time-normalized f0 contours of Mandarin words in conversational speech have been shown to be predictable in part from their contextualized embeddings (C

Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning

TutorialsDGX agent

arXiv:2607.02490v1 Announce Type: new Abstract: Large vision-language models can reason over multimodal inputs by generating textual chains of thought (CoT). A key capability exhibited in CoT reasonin

When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling

ResearchDGX agent

arXiv:2607.01727v1 Announce Type: new Abstract: Synthetic data can be scaled along two routes: Source Expansion (SE), which enlarges the source by adding seed materials or generators, and Fixed-Source

Will Scaling Improve Social Simulation with LLMs?

ResearchDGX agent

arXiv:2607.02464v1 Announce Type: new Abstract: Large Language Model (LLM) social simulations are a promising research method, but they are not yet faithful enough to be adopted widely. In this work,

YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

SafetyDGX agent

arXiv:2601.15588v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained fil

2 Jul 2026

A Task-State Representation for Long-Horizon Mobile GUI Agents

AgentsDGX agent

arXiv:2607.00502v1 Announce Type: new Abstract: While long-horizon mobile GUI agents typically rely on thought-action-observation loops, they struggle to separate persistent task states from transient

A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models

Model ReleasesDGX agent

arXiv:2607.00309v1 Announce Type: cross Abstract: We present a real-time musical interface that converts natural-language scene descriptions into evolving procedural soundscapes. A performer types a p

← Previous
1…2627282930…129
Next →