AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
11 Jun 2026

Fanar-Sadiq: A Multi-Agent Architecture for Grounded Islamic QA

AgentsDGX agent

arXiv:2603.08501v3 Announce Type: replace Abstract: Large language models (LLMs) can answer religious knowledge queries fluently, yet they often hallucinate and misattribute sources, which is especial

Findings of the MAGMaR 2026 Shared Task

ResearchDGX agent

arXiv:2606.12295v1 Announce Type: cross Abstract: This overview paper presents the results of the shared task for the second workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMa

FOCUS: DLLMs Know How to Tame Their Compute Bound

TutorialsDGX agent

arXiv:2601.23278v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (DLLMs) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents

ResearchDGX agent

arXiv:2606.12087v1 Announce Type: new Abstract: Training deep search agents requires verifiable questions whose answers remain unavailable until sufficient evidence has been acquired through search. E

FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

Model ReleasesDGX agent

arXiv:2601.04203v2 Announce Type: replace Abstract: We present FronTalk, a benchmark for front-end code generation that pioneers the study of a unique interaction dynamic: conversational code generati

GraphInfer-Bench: Benchmarking LLM's Inference Capability on Graphs

Model ReleasesDGX agent

arXiv:2606.11562v1 Announce Type: cross Abstract: Graph analysis underlies many applications whose answers cannot be looked up in a single record or retrieved along a path: laundering rings, drug repu

GraspLLM: Towards Zero-Shot Generalization on Text-Attributed Graphs with LLMs

Model ReleasesDGX agent

arXiv:2606.11898v1 Announce Type: new Abstract: Research on Text-Attributed Graphs (TAGs) has gained significant attention recently due to its broad applications across various real-world data scenari

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

ResearchDGX agent

arXiv:2606.11429v1 Announce Type: cross Abstract: Speech foundation models often struggle in low-resource domains due to domain mismatch and data scarcity. We propose Gumbel-BEARD, a domain adaptation

I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System

Model ReleasesDGX agent

arXiv:2606.11875v1 Announce Type: new Abstract: Emotional validation - explicitly acknowledging that a user's feelings make sense - has proven therapeutic value but has received little computational a

Improving Cross-Format Robustness in Language Models with Multi-Format Training

Model ReleasesDGX agent

arXiv:2606.11643v1 Announce Type: new Abstract: Large language models often remain sensitive to answer format: a question solved correctly in one form may fail in another semantically equivalent form.

Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation

ResearchDGX agent

arXiv:2601.07506v2 Announce Type: replace Abstract: While large language models (LLMs) are increasingly used as automatic judges for question answering (QA) and other reference-conditioned evaluation

Kuramoto Attention: Synchronizing Self-Attention on the Torus

ResearchDGX agent

arXiv:2606.11585v1 Announce Type: cross Abstract: We introduce Kuramoto attention, a self-attention layer in which each hidden coordinate is an angle. The layer scores tokens by gated cosine similarit

Language Shapes Mental Health Evaluations in Large Language Models

ResearchDGX agent

arXiv:2603.06910v2 Announce Type: replace Abstract: Multilingual large language models (LLMs) are increasingly used in socially sensitive mental health contexts, including support chatbots, screening,

LatticeBridge: Rare-Event Sequential Inference for Faithful Structured Sequence Synthesis

Model ReleasesDGX agent

arXiv:2606.11203v1 Announce Type: new Abstract: Structured sequence generation often requires a model to satisfy several input-derived constraints in a single output. Standard decoding methods may ass

LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization

Model ReleasesDGX agent

arXiv:2510.23320v2 Announce Type: replace-cross Abstract: We introduce LibriConvo, a synthetic conversational speech corpus for speaker diarization and automatic speech recognition (ASR), built by ins

LifeSentence: Language models can encode human life course trajectories from longitudinal panel data

Model ReleasesDGX agent

arXiv:2606.11220v1 Announce Type: new Abstract: Forecasting human life outcomes is important to gain insights into how individuals attain long and healthy lives. Conventional statistical approaches yi

Lius: Translation Model Based Instructional Lingustic Using Continual Instruction Tuning In Kupang Malay

ResearchDGX agent

arXiv:2606.11786v1 Announce Type: new Abstract: Large Language Models (LLMs) offer new potential for translation tasks but often experience performance degradation when handling low-resource languages

LLMpedia: A Transparent Framework to Materialize an LLM's Encyclopedic Knowledge at Scale

Model ReleasesDGX agent

arXiv:2603.24080v2 Announce Type: replace Abstract: Benchmarks like MMLU suggest flagship language models approach factuality saturation above 90%. LLMpedia shows this picture is incomplete. We materi

M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset

ApplicationsDGX agent

arXiv:2510.23508v3 Announce Type: replace Abstract: Existing real-world datasets for multimodal fact-checking have multiple limitations: they contain few instances, cover on only one or two languages,

Massive Open-Vocabulary Keyword Spotting

ResearchDGX agent

arXiv:2606.11279v1 Announce Type: cross Abstract: Automatic speech recognition systems have been shown to under-perform when it comes to transcribing words rarely seen in the training data, namely spe

Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

AgentsDGX agent

arXiv:2606.12291v1 Announce Type: new Abstract: Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical ju

Measuring language complexity from hierarchical reuse of recurring patterns

ResearchDGX agent

arXiv:2606.11531v1 Announce Type: new Abstract: We introduce the ladderpath index as a measure of language complexity grounded in algorithmic information theory. It counts the minimum steps needed to

Measuring Semantic Progress in Multi-turn Dialogue via Information Gain

SafetyDGX agent

arXiv:2606.12332v1 Announce Type: new Abstract: Evaluating multi-turn dialogue is challenging because quality emerges across turns rather than within individual responses. We focus on a key dimension

Multi-Agent Reasoning with Adaptive Worker Allocation for Stance Detection

Model ReleasesDGX agent

arXiv:2606.11609v1 Announce Type: new Abstract: Stance detection requires identifying an author's position toward a target, often from short-form texts where stance is implicit, indirect, or rhetorica

Neuron-based Personality Trait Induction in Large Language Models

ResearchDGX agent

arXiv:2410.12327v2 Announce Type: replace Abstract: Large language models (LLMs) have become increasingly proficient at simulating various personality traits, an important capability for supporting re

Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills

AgentsDGX agent

arXiv:2606.11897v1 Announce Type: new Abstract: Scientific discovery workflows usually contain and rely heavily on lab notes, where researchers record observations, interpret uncertain results, and pl

On The Effectiveness-Fluency Trade-Off In LLM Conditioning: A Systematic Study

ResearchDGX agent

arXiv:2606.12234v1 Announce Type: new Abstract: Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved t

One Jailbreak, Many Tongues: Learning Language-Insensitive Intention Representations for Multilingual Jailbreak Detection

SafetyDGX agent

arXiv:2606.11202v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in applications for global multilingual users, yet safety training remains concentrated in domina

ProHiFlo: Hierarchical Flow Matching with Functional Guidance for De Novo Protein Generation

ResearchDGX agent

arXiv:2606.11243v1 Announce Type: cross Abstract: De novo protein generation has transformative potential in therapeutic design, enzyme engineering, and synthetic biology. While diffusion-based and fl

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

Model ReleasesDGX agent

arXiv:2606.12250v1 Announce Type: new Abstract: Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical abil

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

SafetyDGX agent

arXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri

SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment

SafetyDGX agent

arXiv:2606.11512v1 Announce Type: new Abstract: Large language models increasingly express uncertainty through natural-language statements, yet these expressions often fail to reflect the model's samp

Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version

SafetyDGX agent

arXiv:2606.11399v1 Announce Type: new Abstract: Large Language Models (LLMs) are deployed across cultural contexts but often reflect homogenized values inherited from training data. Evaluations of cul

Schutzen: Evaluating LLM Safety in Bulgarian and German Contexts

SafetyDGX agent

arXiv:2606.11316v1 Announce Type: new Abstract: Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disr

Semantic Grading of Written Answers in Low-Resource Language Bangla Using a Fine-Tuned Lightweight Language Model

ResearchDGX agent

arXiv:2606.11931v1 Announce Type: new Abstract: Bangla is among the world's most widely spoken languages, yet it remains underserved in educational NLP research. In many remote and rural regions, acce

SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora

Model ReleasesDGX agent

arXiv:2602.10908v2 Announce Type: replace Abstract: We present SoftMatcha 2, an ultra-fast and flexible search algorithm that enables search over trillion-scale natural language corpora in under 0.3 s

SOMA-SQL: Resolving Multi-Source Ambiguity in NL-to-SQL via Synthetic Log and Execution Probing

TutorialsDGX agent

arXiv:2606.11424v1 Announce Type: new Abstract: Natural language interfaces to databases aim to translate user questions into executable SQL, yet remain brittle in real-world settings where questions

StanceNakba Shared Task: Actor and Topic-Aware Stance Detection in Public Discourse

ResearchDGX agent

arXiv:2606.12068v1 Announce Type: new Abstract: We present StanceNakba 2026, a shared task on stance detection in polarized social media discourse related to the Palestinian-Israeli conflict, organize

Steering the Noise: Turning Random Perturbations into Effective Descent for Memory-Efficient LLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2601.04710v2 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) achieves strong performance but is often limited by the memory overhead of backpropagation. Zeroth-order (Z

Teaching Diffusion to Speculate Left-to-Right

Model ReleasesDGX agent

arXiv:2606.11552v1 Announce Type: new Abstract: Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial i

The Language You Ask In: Language-Conditioned Ideological Divergence in LLM Analysis of Contested Political Documents

Model ReleasesDGX agent

arXiv:2601.12164v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as analytical tools across multilingual contexts, yet their outputs may carry systemati

The Long Tail, Not the Front Page: Cold-Start Prediction of Crowd Highlight Salience

ResearchDGX agent

arXiv:2606.11654v1 Announce Type: cross Abstract: A social highlighter's most useful signal -- which passages a crowd of readers marks -- exists only for documents people have already read. Can the ag

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes

AgentsDGX agent

arXiv:2606.11470v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved strong performance across natural language processing tasks, yet reliable reasoning remains an open challenge

UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

SafetyDGX agent

arXiv:2606.11740v1 Announce Type: cross Abstract: We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

SafetyDGX agent

arXiv:2606.11681v1 Announce Type: new Abstract: We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phonem

uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking

ApplicationsDGX agent

arXiv:2606.11945v1 Announce Type: new Abstract: This report describes our participation in SemEval-2026 Task 8 on multi-turn retrieval and question answering. The task evaluates conversational systems

Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery

ResearchDGX agent

arXiv:2602.02726v2 Announce Type: replace-cross Abstract: Large language models (LLMs) encode rich semantic information in their hidden states, yet it remains difficult to understand what information

Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization

Model ReleasesDGX agent

arXiv:2606.12373v1 Announce Type: new Abstract: Reinforcement Learning (RL) with verifiable environments has emerged as a powerful approach for enhancing the reasoning capabilities of Large Language M

VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation

Model ReleasesDGX agent

arXiv:2601.03792v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable proficiency in general medical domains. However, their performance significantly degrades

When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.11906v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic varia

When is Your LLM Steerable?

TutorialsDGX agent

arXiv:2606.11599v1 Announce Type: new Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depen

When More Documents Hurt RAG: Mitigating Vector Search Dilution with Domain-Scoped, Model-Agnostic Retrieval

AgentsDGX agent

arXiv:2606.11350v1 Announce Type: new Abstract: Retrieval-augmented generation degrades when scaled to large, heterogeneous document collections, where dense similarity loses discriminative power, and

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models

ResearchDGX agent

arXiv:2510.01157v4 Announce Type: replace Abstract: Speech language models (SLMs) are systems of systems: independent components that unite to achieve a common goal. Despite their heterogeneous nature

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs

Model ReleasesDGX agent

arXiv:2606.12385v1 Announce Type: new Abstract: Modern LLM training pipelines increasingly rely on other models to generate data, filter corpora, judge outputs, and guide development decisions. These

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

SafetyDGX agent

arXiv:2606.12199v1 Announce Type: cross Abstract: Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute

10 Jun 2026

A Continuous-Time Markov Chain Framework for Insertion Language Models

ResearchDGX agent

arXiv:2606.10199v1 Announce Type: cross Abstract: Insertion Language Models (ILMs) offer several advantages over left-to-right generation and mask-based generation. However, existing formulations of i

A Navigable Manifold of Hypothesized Consciousness-Spectrum States in Language Model Representations

Local AiDGX agent

arXiv:2606.09894v1 Announce Type: cross Abstract: Across contemplative, philosophical, and psychological accounts, human consciousness is often described along a similar spectrum, ranging from reactiv

AI Application Gives Users Real-Time Feedback on the Level of Peace in the Social Media Videos They Watch

ResearchDGX agent

arXiv:2601.05232v3 Announce Type: replace Abstract: Most people now get their news from videos on social media, such as YouTube and Facebook, rather than through curated journalism. 'We become what we

An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs

Model ReleasesDGX agent

arXiv:2603.14463v2 Announce Type: replace Abstract: Adapting Large Language Models (LLMs) to high-stakes vertical domains like insurance presents a significant challenge: scenarios demand strict adher

AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style

ResearchDGX agent

arXiv:2603.11482v2 Announce Type: replace-cross Abstract: Evaluating 'anime-like' voices currently relies on costly subjective judgments, yet no standardized objective metric exists. A key challenge i

← Previous
1…3738394041…129
Next →