AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
30 Jul 2026

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2607.26355v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential

The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text

SafetyDGX agent

arXiv:2607.26309v1 Announce Type: cross Abstract: Estimating causal effects of linguistic properties from observational text is difficult because the same document can contain both the treatment of in

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

Model ReleasesDGX agent

arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

ResearchDGX agent

arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.

Voice Memory for Agentic Speech Recognition

AgentsDGX agent

arXiv:2607.26410v1 Announce Type: new Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md

When Does Span-Guided Detoxification Help? Human Preferences and Evaluator Diagnostics in a Controlled Comparison

ResearchDGX agent

arXiv:2607.26795v1 Announce Type: new Abstract: Span-guided rewriting aims to preserve meaning by localizing edits to annotated harmful spans, but the same constraint can leave harmful intent insuffic

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

Model ReleasesDGX agent

arXiv:2607.26348v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

Model ReleasesDGX agent

arXiv:2607.26555v1 Announce Type: new Abstract: Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplifi

Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms

AgentsDGX agent

arXiv:2607.26497v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluat

WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback

SafetyDGX agent

arXiv:2607.26604v1 Announce Type: new Abstract: Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained in

29 Jul 2026

A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings

ResearchDGX agent

arXiv:2607.25202v1 Announce Type: new Abstract: Conversational entrainment is well-studied in monolingual and written contexts, but remains underexplored in spoken code-switching (CSW). We present a n

A scaling law of contextual persistence in human language

ResearchDGX agent

arXiv:2607.25184v1 Announce Type: new Abstract: Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we sh

A Study of Crosslinguistic Influence in Language Models

ResearchDGX agent

arXiv:2601.21587v2 Announce Type: replace Abstract: The sequential acquisition of languages inevitably leads to Crosslinguistic Influence (CLI), where the syntactic properties of a first language (L1)

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

Model ReleasesDGX agent

arXiv:2607.25881v1 Announce Type: new Abstract: We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were indepen

An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data

ResearchDGX agent

arXiv:2503.07303v3 Announce Type: replace Abstract: Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulai

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

ApplicationsDGX agent

arXiv:2607.25852v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference without changing the target distribution, but no single drafting structure performs best

Constrained CTC Decoding for Efficient Diacritic Restoration

ResearchDGX agent

arXiv:2607.18946v2 Announce Type: replace Abstract: In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling f

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

ResearchDGX agent

arXiv:2607.25291v1 Announce Type: new Abstract: The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical

Deep Label-Wise Attentive Temporal Convolutional Networks Improve Medical Coding

TutorialsDGX agent

arXiv:2607.25129v1 Announce Type: new Abstract: Medical coding is the task of assigning a set of diagnosis and procedure codes for a hospitalization using recorded notes. It requires aggregating infor

Evaluation of Adversarial Robustness in Arabic Language Models

SafetyDGX agent

arXiv:2607.25814v1 Announce Type: new Abstract: The emergence of the recent outstanding capabilities of Arabic Language Models has opened doors for exposing their vulnerabilities. One of the major sec

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

SafetyDGX agent

arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alig

Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Methods

ApplicationsDGX agent

arXiv:2502.11150v5 Announce Type: replace Abstract: Automatic methods for scoring text readability have been studied for over a century, and are widely used in research and in user-facing applications

FlashEvaluator: Expanding Search Space with Parallel Sequence-Level Evaluation

Local AiDGX agent

arXiv:2603.02565v2 Announce Type: replace-cross Abstract: The Generator-Evaluator (G-E) framework generates K candidate sequences and uses an evaluator to select the highest-scoring one, which is wide

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

Model ReleasesDGX agent

arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and reposito

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

SafetyDGX agent

arXiv:2508.05775v3 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language

Influence of Prompt Engineering on Small Language Models for Guarded Query Routing

Model ReleasesDGX agent

arXiv:2607.24801v1 Announce Type: cross Abstract: We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

Model ReleasesDGX agent

arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recogniz

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

Model ReleasesDGX agent

arXiv:2607.25642v1 Announce Type: cross Abstract: Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

Model ReleasesDGX agent

arXiv:2607.26015v1 Announce Type: new Abstract: Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented featu

Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization

Local AiDGX agent

arXiv:2607.25228v1 Announce Type: new Abstract: Column annotation (CA), including column type annotation (CTA) and column property annotation (CPA), aims to identify the meanings of table columns and

Language as a Material Interface for Creative LLM Interaction

ResearchDGX agent

arXiv:2607.24753v1 Announce Type: cross Abstract: Although directive prompting is the predominant way to interact with Large Language Models (LLMs), many creative practices rely on language that is op

M^2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

Model ReleasesDGX agent

arXiv:2510.13434v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by m

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Model ReleasesDGX agent

arXiv:2607.24904v1 Announce Type: cross Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streamin

Med-R^3: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning

Model ReleasesDGX agent

arXiv:2507.23541v5 Announce Type: replace Abstract: In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despi

Memory for Large Language Models

Model ReleasesDGX agent

arXiv:2607.25380v1 Announce Type: new Abstract: Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Model ReleasesDGX agent

arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastro

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

Model ReleasesDGX agent

arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, mu

Neuromorphic Diffusion Language Models: Addressing Compute and Memory Bottlenecks via Sparsity and Block Denoising

Model ReleasesDGX agent

arXiv:2607.24841v1 Announce Type: new Abstract: Autoregressive (AR) large language models (LLMs) are inherently inefficient at inference time because each generated token requires accessing the full s

Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance

SafetyDGX agent

arXiv:2607.25507v1 Announce Type: new Abstract: Transformer language models are usually analyzed through vector geometry, yet ordered context and rotary position encoding introduce explicit phase stru

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers

ApplicationsDGX agent

arXiv:2512.17351v2 Announce Type: replace Abstract: Understanding architectural differences in language models is challenging, especially at academic-scale pretraining (e.g., 1.3B parameters, 100B tok

PILA: Plug-and-Play Insertion for LLM-native Advertising

TutorialsDGX agent

arXiv:2607.25590v1 Announce Type: new Abstract: How to monetize large language models (LLMs) by naturally integrating sponsored content into their responses, known as LLM-native advertising, has recen

PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning

Model ReleasesDGX agent

arXiv:2508.00344v5 Announce Type: replace Abstract: Large Language Models (LLMs) have shown remarkable advancements in tackling agent-oriented tasks. Despite their potential, existing work faces chall

Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections

Model ReleasesDGX agent

arXiv:2607.25953v1 Announce Type: new Abstract: As LLMs increasingly mediate the political information citizens rely on, there is still no standardized way to assess whether they do so responsibly. We

Ranked by Position: Order Sensitivity as an Exploitable Attack Surface in LLM Listwise Recommenders

SafetyDGX agent

arXiv:2607.24869v1 Announce Type: cross Abstract: Large language models (LLMs) used as listwise rerankers in recommendation systems suffer from position bias when serializing candidate sets into promp

Research Report on Noise-Shaped One-Bit Coefficients in Discrete Polynomial Fourier Extension

Model ReleasesDGX agent

arXiv:2607.24868v1 Announce Type: new Abstract: This report studies noise-shaped one-bit coefficients in normalized discrete polynomial Fourier extension. For first-order Sigma-Delta quantization, the

Retrieval, not hallucinations, will be the limiting factor for LLM-based clinical AI tools

ResearchDGX agent

arXiv:2607.24793v1 Announce Type: cross Abstract: Discussions around large language model (LLM) errors in clinical artificial intelligence (AI) generally center around precision errors like hallucinat

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

Model ReleasesDGX agent

arXiv:2607.25886v1 Announce Type: cross Abstract: Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capa

SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration

Model ReleasesDGX agent

arXiv:2606.18902v2 Announce Type: replace Abstract: Context engineering has emerged as a primary lever for improving AI systems without parameter updates. Recent work showing that textual gradients do

SciClaimSeekers at CheckThat! 2026: Retrieving Scientific Sources for Social Media Claims with LLM Reranking

ResearchDGX agent

arXiv:2607.24803v1 Announce Type: cross Abstract: Scientific claims often spread on social media faster than they can be verified, while posts rarely link to the original scholarly sources. To tackle

Shieldstral

Model ReleasesDGX agent

arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s

SourceMinds at CheckThat! 2026: NLI-Grounded Citation Auditing in a Multi-Agent Pipeline for Full Fact-Checking Article Generation

AgentsDGX agent

arXiv:2607.24802v1 Announce Type: cross Abstract: This paper presents our system for Task 3 of the CLEF 2026 CheckThat! Lab, which focuses on generating full fact-checking articles from claims, veraci

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

Model ReleasesDGX agent

arXiv:2607.25716v1 Announce Type: new Abstract: Federated learning (FL) enables privacy-preserving training of automatic speech recognition (ASR) systems across distributed data sources, yet its appli

Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control

TutorialsDGX agent

arXiv:2607.25337v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them

TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

Model ReleasesDGX agent

arXiv:2607.24750v1 Announce Type: new Abstract: Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliabl

Toward a systematic method for identifying language areas

ResearchDGX agent

arXiv:2607.25305v1 Announce Type: new Abstract: Macroareas are geographical areas used in typological research for grouping variables of interest. In linguistic typology, languages in a given macroare

UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams

Model ReleasesDGX agent

arXiv:2607.26017v1 Announce Type: new Abstract: Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over bound

VisRAG2.0: Mitigating Visual Hallucinations via Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2510.09733v2 Announce Type: replace Abstract: Visual Retrieval-Augmented Generation (VRAG) has emerged as a promising paradigm for equipping Vision-Language Models (VLMs) with external visual ev

VisualPatchWorld: Code World Models as Latent Structured Representations for Planning

TutorialsDGX agent

arXiv:2607.25236v1 Announce Type: new Abstract: Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a for

When Algorithms Meet Artists: Semantic Compression and Stake-holder Marginalisation in Public AI-Art Discourse (2013-2025)

ApplicationsDGX agent

arXiv:2508.03037v5 Announce Type: replace Abstract: Artists occupy a paradoxical position in generative AI. Their own work trains models that now compete with them, replicate their styles, and reshape

← Previous
1…1516171819…128
Next →