AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
7 Aug 2026

Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages

Model ReleasesDGX agent

arXiv:2608.05163v1 Announce Type: new Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test th

Who Checks the Citations? Benchmarking Legal Hallucination Detection

Model ReleasesDGX agent

arXiv:2606.21155v2 Announce Type: replace Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictio

Zero-Shot Multi-Disease Labeling of Chest, Abdomen, and Pelvis CT Reports Using Open-Weight Large Language Models: The Effect of Labeling Conventions


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2506.03259v3 Announce Type: replace Abstract: Purpose: To compare five lightweight open-weight large language models (LLMs) with a rule-based algorithm (RBA) and fine-tuned RadBERT for zero-shot

6 Aug 2026

A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy

ResearchDGX agent

arXiv:2608.04808v1 Announce Type: new Abstract: Part-of-speech tagging for low-resource languages remains challenging due to limited annotated data, especially for linguistically complex languages. Ga

Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems

Model ReleasesDGX agent

arXiv:2608.04746v1 Announce Type: new Abstract: LLM agents that persist across sessions accumulate stored memories whose validity varies enormously by content type, yet existing memory architectures t

Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study

Model ReleasesDGX agent

arXiv:2608.04576v1 Announce Type: new Abstract: Humanitarian reports are long, noisy, and multi-topic, making it difficult to consolidate decision-relevant causal evidence. We present a ReliefWeb stud

DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

SafetyDGX agent

arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific

DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models

ResearchDGX agent

arXiv:2608.04477v1 Announce Type: cross Abstract: Cloud-based language model services routinely process prompts containing sensitive information. Obfuscation-based defenses---including ObfusLM, Sentin

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

Model ReleasesDGX agent

arXiv:2608.05004v1 Announce Type: new Abstract: Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including 'delusi

Do Language Models Know Their Slang? Queer Slang Understanding in User-Generated Content

ResearchDGX agent

arXiv:2608.04847v1 Announce Type: new Abstract: Despite its cultural relevance and diffusion, queer slang remains underrepresented in Natural Language Processing research. Towards addressing this gap,

Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

ResearchDGX agent

arXiv:2604.00209v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private

Document Optimization for Black-Box Retrieval via Reinforcement Learning

ResearchDGX agent

arXiv:2604.05087v3 Announce Type: replace Abstract: Document expansion is a classical technique for improving retrieval quality, and is attractive since it shifts computation offline, avoiding additio

Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?

ResearchDGX agent

arXiv:2608.04928v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning offers a window into the decision-making of large language models (LLMs), which can be monitored for target behaviors b

EdgeLM: Edge Demonstrations for Language Models' Table Understanding

ResearchDGX agent

arXiv:2608.04390v1 Announce Type: new Abstract: Large language models (LLMs) perform table-centric prediction through in-context learning, making demonstration selection critical to performance. Exist

Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks

Model ReleasesDGX agent

arXiv:2608.04286v1 Announce Type: new Abstract: Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot

AgentsDGX agent

arXiv:2608.04709v1 Announce Type: new Abstract: This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation

Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs

Model ReleasesDGX agent

arXiv:2608.04488v1 Announce Type: new Abstract: Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high

Enhancing Trustworthy Clinical Diagnosis Decision-Making in Large Language Models via Etiology-Aware Attention Supervision

Model ReleasesDGX agent

arXiv:2508.00285v2 Announce Type: replace Abstract: Objective: Large Language Models (LLMs) have demonstrated strong capabilities in medical text understanding and generation. However, their trustwort

Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

SafetyDGX agent

arXiv:2608.04885v1 Announce Type: cross Abstract: Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validat

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

ResearchDGX agent

arXiv:2608.04646v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

ResearchDGX agent

arXiv:2608.04899v1 Announce Type: new Abstract: Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as

GEB-Bench: Abstract Structures Told in Many Voices

Model ReleasesDGX agent

arXiv:2608.04111v1 Announce Type: cross Abstract: Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstr

German parties shifted towards intuition-based rhetoric after the far right's parliamentary breakthrough

ResearchDGX agent

arXiv:2608.05075v1 Announce Type: new Abstract: The spread of misinformation is widely perceived as a threat to democratic deliberation, yet how political elites' rhetorical commitments to truth shift

IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

Model ReleasesDGX agent

arXiv:2608.04703v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where an

K-EXAONE 2.0 Technical Report

Model ReleasesDGX agent

arXiv:2608.04505v1 Announce Type: new Abstract: This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward glo

Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention

Model ReleasesDGX agent

arXiv:2608.04678v1 Announce Type: new Abstract: Papers 1-2 of the Kathleen series showed that a byte-level, attention-free architecture built from a wavetable encoder and multi-scale reverberant state

Language Models Generalize to Human-like Word Order Preferences

SafetyDGX agent

arXiv:2608.05028v1 Announce Type: new Abstract: A central question in language acquisition is whether linguistic biases can emerge from general learning mechanisms operating over underdetermined input

Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language

Model ReleasesDGX agent

arXiv:2608.04186v1 Announce Type: new Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs).

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

AgentsDGX agent

arXiv:2608.04056v1 Announce Type: new Abstract: When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP sys

Leveraging Machine Learning to Gain Insights on Quantum Thermodynamic Entropy

ResearchDGX agent

arXiv:2305.06177v1 Announce Type: cross Abstract: We present a thermodynamic analysis of a quantum engine that uses a single quantum particle as its working fluid, inspired by Szilard's classical sing

LLM-based Vulnerability Discovery in Business Process Documentation

ResearchDGX agent

arXiv:2608.04271v1 Announce Type: cross Abstract: Just like software and hardware, business processes are susceptible to vulnerabilities that can lead to product quality issues, delays, and increased

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

SafetyDGX agent

arXiv:2606.18709v2 Announce Type: replace Abstract: Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item mea

MathDebugger: Detecting and Diagnosing Errors in Synthetic Mathematical Data

Model ReleasesDGX agent

arXiv:2502.19058v2 Announce Type: replace Abstract: Synthetic mathematical data has become an important resource for scaling the reasoning capabilities of large language models, yet errors in generate

Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap

Model ReleasesDGX agent

arXiv:2608.04160v1 Announce Type: new Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so th

Neighborhood-Aware Dual Biomedical Entity Linking

ResearchDGX agent

arXiv:2608.04144v1 Announce Type: cross Abstract: Biomedical entity linking grounds mentions in clinical and scientific text to entities in a curated knowledge base (KB) with ontological structure, wh

Neurocomputational Mechanisms of Syntactic Transfer in Bilingual Sentence Production

ApplicationsDGX agent

arXiv:2601.18056v2 Announce Type: replace Abstract: We discuss the benefits of incorporating oscillatory neural mechanisms into the study of bilingual production errors and their traditionally documen

NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap

Model ReleasesDGX agent

arXiv:2608.04397v1 Announce Type: new Abstract: We introduce NOLLI, a procedurally generated English-Korean puzzle benchmark designed to diagnose where Korean performance gaps arise. It comprises 15 p

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

Model ReleasesDGX agent

arXiv:2606.17188v3 Announce Type: replace-cross Abstract: Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking b

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance

SafetyDGX agent

arXiv:2608.04524v1 Announce Type: new Abstract: Synthetic generation of Cognitive Behavioral Therapy (CBT) sessions is challenged by two competing demands: adhering to strict therapeutic structure whi

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning

SafetyDGX agent

arXiv:2608.05080v1 Announce Type: cross Abstract: Critic-free group-based reinforcement learning has become a scalable approach for post-training large language models. However, most existing methods

Predict, Then Retrieve: Cross-Instance Future-State Retrieval from Video Prefixes

Model ReleasesDGX agent

arXiv:2608.04426v1 Announce Type: cross Abstract: We introduce Predictive State Retrieval (PSR), a task in which a model observes a short video prefix and a temporal question about an object's future

Preverbal Uninflected and Underived Roots in Mapudungun. Wuno and Its Implications

ResearchDGX agent

arXiv:2608.04869v1 Announce Type: new Abstract: This study examines the grammatical status of preverbal uninflected and underived roots in Mapudungun, with particular focus on wuno 'return/re-'. Throu

Reachability in 3-VAS

ResearchDGX agent

arXiv:2608.04786v1 Announce Type: new Abstract: We settle the exact complexity of the reachability problem in (stateless) vector addition systems (VAS) in fixed low dimension. In dimensions 2-4 it has

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

Model ReleasesDGX agent

arXiv:2608.04939v1 Announce Type: new Abstract: Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

ResearchDGX agent

arXiv:2608.05148v1 Announce Type: new Abstract: Procedural generators produce useful verifiable reasoning problems at scale, but have received less attention as data for completion-supervised fine-tun

Relational Response Fields: A General Theory of Black-Box LLM Response Consistency and Recovery

ResearchDGX agent

arXiv:2608.04552v1 Announce Type: new Abstract: Black-box language-model reliability is commonly pursued by sampling, prompting, voting, verifying, or iteratively revising individual answers. We ask a

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

Model ReleasesDGX agent

arXiv:2608.04569v1 Announce Type: new Abstract: Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring unit

Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling

ResearchDGX agent

arXiv:2608.04554v1 Announce Type: new Abstract: Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available.

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Model ReleasesDGX agent

arXiv:2608.04514v1 Announce Type: new Abstract: Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-co

Right Reset: Chunking by Prefix Removal

ResearchDGX agent

arXiv:2608.04330v1 Announce Type: new Abstract: Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right-hand tokens w

RingSQL: Schema-Independent Synthetic Data Generation for Text-to-SQL Reinforcement Learning

ResearchDGX agent

arXiv:2601.05451v2 Announce Type: replace-cross Abstract: Recent advances in text-to-SQL have been driven by larger models, better datasets, and new training methods like RLVR. However, progress remai

Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?

Model ReleasesDGX agent

arXiv:2608.05097v1 Announce Type: new Abstract: Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same

Searching for Sound-Meaning Collisions: Graph-Based Affordance Retrieval and Multi-Evaluator Ranking for Pun Translation at CLEF 2026 JOKER Task 2

ResearchDGX agent

arXiv:2608.04299v1 Announce Type: new Abstract: Fifteen years ago, Low proposed that pun translators should stop searching for equivalent words and instead search for new points of contact between sou

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.04244v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks

Simile Understanding in Text-to-Image Models: An Evaluation Framework

ResearchDGX agent

arXiv:2608.04750v1 Announce Type: cross Abstract: Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

Model ReleasesDGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

Social Pressure Breaks Majority Voting in LLM Safety Panels

SafetyDGX agent

arXiv:2608.04415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization

Model ReleasesDGX agent

arXiv:2608.04084v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) networks pursue specialization through learned routers, gates, and load-balancing losses, yet at matched total-parameter budg

SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts

SafetyDGX agent

arXiv:2608.04962v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves the reasoning capabilities of large language models, but autoregressive rollout generation remains

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

AgentsDGX agent

arXiv:2608.05126v1 Announce Type: new Abstract: Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interac

← Previous
1…45678…128
Next →