AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Model Releases

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

DGX agent

arXiv:2607.26355v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential

model-releasesarxiv-cs-cl
30 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text

DGX agent

arXiv:2607.26309v1 Announce Type: cross Abstract: Estimating causal effects of linguistic properties from observational text is difficult because the same document can contain both the treatment of in

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

DGX agent

arXiv:2503.10647v2 Announce Type: replace Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dim

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

DGX agent

arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once

model-releasesarxiv-cs-cl
30 Jul 2026
Research

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

DGX agent

arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.

researcharxiv-cs-cl
30 Jul 2026
Agents

Voice Memory for Agentic Speech Recognition

DGX agent

arXiv:2607.26410v1 Announce Type: new Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md

agentsarxiv-cs-cl
30 Jul 2026
Research

When Does Span-Guided Detoxification Help? Human Preferences and Evaluator Diagnostics in a Controlled Comparison

DGX agent

arXiv:2607.26795v1 Announce Type: new Abstract: Span-guided rewriting aims to preserve meaning by localizing edits to annotated harmful spans, but the same constraint can leave harmful intent insuffic

researcharxiv-cs-cl
30 Jul 2026
Model Releases

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

DGX agent

arXiv:2607.26348v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Where Detectors Fail: Closing the Tail-Domain Gap with Expert-Guided Mutual Distillation

DGX agent

arXiv:2607.26555v1 Announce Type: new Abstract: Multimodal fake news detectors often generalize poorly across domains because they learn to trust unreliable evidence: domain-specific shortcuts amplifi

model-releasesarxiv-cs-cl
30 Jul 2026
Agents

Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms

DGX agent

arXiv:2607.26497v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluat

agentsarxiv-cs-cl
30 Jul 2026
Safety

WikiLoop: Jointly Learning to Build and Navigate Agent-Native Wikis with Downstream Feedback

DGX agent

arXiv:2607.26604v1 Announce Type: new Abstract: Knowledge-base construction and querying are typically optimized in isolation: retrieval-augmented agents operate over a fixed, externally maintained in

safetyarxiv-cs-cl
30 Jul 2026
Research

A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings

DGX agent

arXiv:2607.25202v1 Announce Type: new Abstract: Conversational entrainment is well-studied in monolingual and written contexts, but remains underexplored in spoken code-switching (CSW). We present a n

researcharxiv-cs-cl
29 Jul 2026
Research

A scaling law of contextual persistence in human language

DGX agent

arXiv:2607.25184v1 Announce Type: new Abstract: Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we sh

researcharxiv-cs-cl
29 Jul 2026
Research

A Study of Crosslinguistic Influence in Language Models

DGX agent

arXiv:2601.21587v2 Announce Type: replace Abstract: The sequential acquisition of languages inevitably leads to Crosslinguistic Influence (CLI), where the syntactic properties of a first language (L1)

researcharxiv-cs-cl
29 Jul 2026
Model Releases

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation

DGX agent

arXiv:2607.25881v1 Announce Type: new Abstract: We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were indepen

model-releasesarxiv-cs-cl
29 Jul 2026
Research

An Information-Theoretic Approach to Identifying Formulaic Clusters in Textual Data

DGX agent

arXiv:2503.07303v3 Announce Type: replace Abstract: Texts, whether literary or historical, exhibit structural and stylistic patterns shaped by their purpose, authorship, and cultural context. Formulai

researcharxiv-cs-cl
29 Jul 2026
Applications

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

DGX agent

arXiv:2607.25852v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference without changing the target distribution, but no single drafting structure performs best

applicationsarxiv-cs-cl
29 Jul 2026
Research

Constrained CTC Decoding for Efficient Diacritic Restoration

DGX agent

arXiv:2607.18946v2 Announce Type: replace Abstract: In this work, we address diacritic restoration for Arabic speech transcripts. Most speech data are undiacritized, limiting the ability of modeling f

researcharxiv-cs-cl
29 Jul 2026
Research

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

DGX agent

arXiv:2607.25291v1 Announce Type: new Abstract: The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical

researcharxiv-cs-cl
29 Jul 2026
Tutorials

Deep Label-Wise Attentive Temporal Convolutional Networks Improve Medical Coding

DGX agent

arXiv:2607.25129v1 Announce Type: new Abstract: Medical coding is the task of assigning a set of diagnosis and procedure codes for a hospitalization using recorded notes. It requires aggregating infor

tutorialsarxiv-cs-cl
29 Jul 2026
Safety

Evaluation of Adversarial Robustness in Arabic Language Models

DGX agent

arXiv:2607.25814v1 Announce Type: new Abstract: The emergence of the recent outstanding capabilities of Arabic Language Models has opened doors for exposing their vulnerabilities. One of the major sec

safetyarxiv-cs-cl
29 Jul 2026
Safety

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

DGX agent

arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alig

safetyarxiv-cs-cl
29 Jul 2026
Applications

Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Methods

DGX agent

arXiv:2502.11150v5 Announce Type: replace Abstract: Automatic methods for scoring text readability have been studied for over a century, and are widely used in research and in user-facing applications

applicationsarxiv-cs-cl
29 Jul 2026
Local Ai

FlashEvaluator: Expanding Search Space with Parallel Sequence-Level Evaluation

DGX agent

arXiv:2603.02565v2 Announce Type: replace-cross Abstract: The Generator-Evaluator (G-E) framework generates K candidate sequences and uses an evaluator to select the highest-scoring one, which is wide

local-aiarxiv-cs-cl
29 Jul 2026
Model Releases

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

DGX agent

arXiv:2607.25589v1 Announce Type: cross Abstract: Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and reposito

model-releasesarxiv-cs-cl
29 Jul 2026
Safety

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

DGX agent

arXiv:2508.05775v3 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language

safetyarxiv-cs-cl
29 Jul 2026
Model Releases

Influence of Prompt Engineering on Small Language Models for Guarded Query Routing

DGX agent

arXiv:2607.24801v1 Announce Type: cross Abstract: We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context

DGX agent

arXiv:2607.25375v1 Announce Type: new Abstract: India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recogniz

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

DGX agent

arXiv:2607.25642v1 Announce Type: cross Abstract: Instruction-based Image Editing (IIE) aims to transform a given image into a new one based on textual instructions. Advances in Large Language Models

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do

DGX agent

arXiv:2607.26015v1 Announce Type: new Abstract: Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented featu

model-releasesarxiv-cs-cl
29 Jul 2026
Local Ai

Interpretable Column Annotation with LLM-Symbolized Decision Process Materialization

DGX agent

arXiv:2607.25228v1 Announce Type: new Abstract: Column annotation (CA), including column type annotation (CTA) and column property annotation (CPA), aims to identify the meanings of table columns and

local-aiarxiv-cs-cl
29 Jul 2026
Research

Language as a Material Interface for Creative LLM Interaction

DGX agent

arXiv:2607.24753v1 Announce Type: cross Abstract: Although directive prompting is the predominant way to interact with Large Language Models (LLMs), many creative practices rely on language that is op

researcharxiv-cs-cl
29 Jul 2026
Model Releases

M^2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

DGX agent

arXiv:2510.13434v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by m

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

DGX agent

arXiv:2607.24904v1 Announce Type: cross Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streamin

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Med-R^3: Enhancing Medical Retrieval-Augmented Reasoning of LLMs via Progressive Reinforcement Learning

DGX agent

arXiv:2507.23541v5 Announce Type: replace Abstract: In medical scenarios, effectively retrieving external knowledge and leveraging it for rigorous logical reasoning is of significant importance. Despi

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Memory for Large Language Models

DGX agent

arXiv:2607.25380v1 Announce Type: new Abstract: Memory has evolved into a foundational architectural dimension in large language models (LLMs), shifting from an implicit byproduct of computation to a

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

DGX agent

arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastro

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

DGX agent

arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, mu

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Neuromorphic Diffusion Language Models: Addressing Compute and Memory Bottlenecks via Sparsity and Block Denoising

DGX agent

arXiv:2607.24841v1 Announce Type: new Abstract: Autoregressive (AR) large language models (LLMs) are inherently inefficient at inference time because each generated token requires accessing the full s

model-releasesarxiv-cs-cl
29 Jul 2026
Safety

Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance

DGX agent

arXiv:2607.25507v1 Announce Type: new Abstract: Transformer language models are usually analyzed through vector geometry, yet ordered context and rotary position encoding introduce explicit phase stru

safetyarxiv-cs-cl
29 Jul 2026
Applications

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers

DGX agent

arXiv:2512.17351v2 Announce Type: replace Abstract: Understanding architectural differences in language models is challenging, especially at academic-scale pretraining (e.g., 1.3B parameters, 100B tok

applicationsarxiv-cs-cl
29 Jul 2026
Tutorials

PILA: Plug-and-Play Insertion for LLM-native Advertising

DGX agent

arXiv:2607.25590v1 Announce Type: new Abstract: How to monetize large language models (LLMs) by naturally integrating sponsored content into their responses, known as LLM-native advertising, has recen

tutorialsarxiv-cs-cl
29 Jul 2026
Model Releases

PilotRL: Training Language Model Agents via Global Planning-Guided Progressive Reinforcement Learning

DGX agent

arXiv:2508.00344v5 Announce Type: replace Abstract: Large Language Models (LLMs) have shown remarkable advancements in tackling agent-oriented tasks. Despite their potential, existing work faces chall

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections

DGX agent

arXiv:2607.25953v1 Announce Type: new Abstract: As LLMs increasingly mediate the political information citizens rely on, there is still no standardized way to assess whether they do so responsibly. We

model-releasesarxiv-cs-cl
29 Jul 2026
Safety

Ranked by Position: Order Sensitivity as an Exploitable Attack Surface in LLM Listwise Recommenders

DGX agent

arXiv:2607.24869v1 Announce Type: cross Abstract: Large language models (LLMs) used as listwise rerankers in recommendation systems suffer from position bias when serializing candidate sets into promp

safetyarxiv-cs-cl
29 Jul 2026
Model Releases

Research Report on Noise-Shaped One-Bit Coefficients in Discrete Polynomial Fourier Extension

DGX agent

arXiv:2607.24868v1 Announce Type: new Abstract: This report studies noise-shaped one-bit coefficients in normalized discrete polynomial Fourier extension. For first-order Sigma-Delta quantization, the

model-releasesarxiv-cs-cl
29 Jul 2026
Research

Retrieval, not hallucinations, will be the limiting factor for LLM-based clinical AI tools

DGX agent

arXiv:2607.24793v1 Announce Type: cross Abstract: Discussions around large language model (LLM) errors in clinical artificial intelligence (AI) generally center around precision errors like hallucinat

researcharxiv-cs-cl
29 Jul 2026
Model Releases

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

DGX agent

arXiv:2607.25886v1 Announce Type: cross Abstract: Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capa

model-releasesarxiv-cs-cl
29 Jul 2026
← Previous
1…1920212223…160
Next →