AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
2 Jun 2026

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models

Model ReleasesDGX agent

arXiv:2606.02041v1 Announce Type: new Abstract: Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate.

SindBERT, the Sailor: Charting the Seas of Turkish NLP

Model ReleasesDGX agent

arXiv:2510.21364v2 Announce Type: replace Abstract: Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. Wit

SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

Model ReleasesDGX agent

arXiv:2606.02540v1 Announce Type: new Abstract: Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party ski


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

Model ReleasesDGX agent

arXiv:2603.08000v2 Announce Type: replace Abstract: Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasonin

SN-WER: Script-Normalized WER for Multi-Script Indic ASR Evaluation

ResearchDGX agent

arXiv:2606.02548v1 Announce Type: new Abstract: Word Error Rate (WER) is the dominant metric for automatic speech recognition (ASR), but it can overestimate errors when references and hypotheses encod

Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

ResearchDGX agent

arXiv:2606.01479v1 Announce Type: new Abstract: Integrating large language models (LLMs) into text-to-speech (TTS) systems has improved speech expressiveness, yet interpretable emotional control remai

Stabilizing Policy Optimization via Logits Convexity

SafetyDGX agent

arXiv:2603.00963v2 Announce Type: replace-cross Abstract: While reinforcement learning (RL) has been central to the recent success of large language models (LLMs), RL optimization is notoriously unsta

Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction

Model ReleasesDGX agent

arXiv:2601.20803v2 Announce Type: replace Abstract: This paper presents several strategies to automatically obtain additional examples for in-context learning, effectively transforming relation extrac

SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

AgentsDGX agent

arXiv:2602.23866v2 Announce Type: replace-cross Abstract: Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL). However, RL training

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

SafetyDGX agent

arXiv:2606.00851v1 Announce Type: cross Abstract: Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or a

TalkTag: Fine-Grained Morphosyntactic Error Annotation for Transcribed Speech

ResearchDGX agent

arXiv:2606.01820v1 Announce Type: new Abstract: Fine-grained morphosyntactic error annotation is important in clinical and developmental language research, yet it is labour-intensive, expert-dependent

Task Structure Reverses Layerwise State Encoding in Sequence Models

ResearchDGX agent

arXiv:2606.00926v1 Announce Type: cross Abstract: Mechanistic studies of sequence models often treat layerwise state encodings as architectural traits: recurrent models concentrate readable state, att

The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete

Model ReleasesDGX agent

arXiv:2606.00048v1 Announce Type: cross Abstract: Prior research has established that instruction-tuned large language models exhibit left-of-center political bias, measured exclusively through abstra

The Social Cost of Intelligence: Emergence, Propagation, and Amplification of Stereotypical Bias in Multi-Agent Systems

SafetyDGX agent

arXiv:2510.10943v2 Announce Type: replace-cross Abstract: Bias in large language models (LLMs) remains a persistent challenge, often leading to stereotyping and unfair treatment across social groups.

Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs

ResearchDGX agent

arXiv:2606.01168v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has significantly enhanced LLM reasoning, yet often incurs substantial computational overhead due to 'overthinking': generating e

ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind

TutorialsDGX agent

arXiv:2505.22961v3 Announce Type: replace Abstract: Large language models (LLMs) have shown promising potential in persuasion, but existing works on training LLM persuaders are still preliminary. Nota

Toward Responsible and Epistemically Grounded Multilingual LLMs for Computational Social Science and Humanities

SafetyDGX agent

arXiv:2606.00596v1 Announce Type: new Abstract: Large language models have rapidly evolved in multilingual competence and reasoning capacity, enabling their integration into Social Sciences and Humani

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

Model ReleasesDGX agent

arXiv:2606.00919v1 Announce Type: new Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - re

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

Model ReleasesDGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

Training Prompt Matters: State-Adaptive Optimization for Robust Fine-Tuning

ResearchDGX agent

arXiv:2606.01967v1 Announce Type: new Abstract: While prompt engineering is instrumental in maximizing the capabilities of Large Language Models (LLMs) during inference, the role of prompts during tra

Transferable Self-Harm Surveillance from Emergency Department Triage Notes Using an Evidence-Augmented Machine Learning Approach

ResearchDGX agent

arXiv:2606.02545v1 Announce Type: new Abstract: Self-harm is a major public health concern, but current surveillance relying on hospital presentations is inadequate due to the low sensitivity of diagn

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher

TutorialsDGX agent

arXiv:2606.01000v1 Announce Type: cross Abstract: Weak-to-strong generalization studies how to improve a strong student using supervision from a weaker teacher when reliable labels are scarce. We view

Trust Region On-Policy Distillation

SafetyDGX agent

arXiv:2606.01249v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

Model ReleasesDGX agent

arXiv:2606.01456v1 Announce Type: cross Abstract: Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sale

TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation

Model ReleasesDGX agent

arXiv:2606.02320v1 Announce Type: new Abstract: Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmar

Uncovering Temporal Framing in the News

ResearchDGX agent

arXiv:2606.00294v1 Announce Type: new Abstract: Temporal language does more than place events on a timeline. In news discourse, references to the past, present, and future can function as rhetorical d

UniD^3: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning

Model ReleasesDGX agent

arXiv:2606.01394v1 Announce Type: new Abstract: Systematic characterization of drug-disease relationships is essential for drug discovery and repurposing, yet is hindered by the heterogeneity and rapi

Unified Context Evolution for LLM Agents

AgentsDGX agent

arXiv:2606.02304v1 Announce Type: new Abstract: LLM-based agents can solve multi-step interactive tasks by combining reasoning with environment feedback, yet each episode starts from the same fixed co

Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention

Model ReleasesDGX agent

arXiv:2606.01243v1 Announce Type: new Abstract: Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over ex

Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning

ResearchDGX agent

arXiv:2606.02020v1 Announce Type: new Abstract: This paper investigates the entropy dynamics of Chain-of-Thought (CoT) and uncovers a consistent two-phase structure: an Uncertainty Region of explorati

VERA: Variational Inference Framework for Jailbreaking Large Language Models

SafetyDGX agent

arXiv:2506.22666v3 Announce Type: replace-cross Abstract: The rise of API-only access to state-of-the-art LLMs highlights the need for effective black-box jailbreak methods to identify model vulnerabi

WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

Model ReleasesDGX agent

arXiv:2510.22276v3 Announce Type: replace-cross Abstract: Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing Engl

WAXAL-NET: Finetuned Edge ASR Across 19 African Languages

ResearchDGX agent

arXiv:2606.02375v1 Announce Type: new Abstract: We evaluate whether compact domain-specialized ASR models can outperform massively multilingual foundation models for conversational African speech acro

What to Format and How: A Benchmark and Workflow Approach for Document Formatting

Model ReleasesDGX agent

arXiv:2606.01936v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have opened up new possibilities for automated document formatting. However, real-world formatting often

When Is 0.1% Enough? Analyzing the Combined Effects of Dimensionality Reduction and Quantization on Text Embedding Compression

ResearchDGX agent

arXiv:2606.01074v1 Announce Type: new Abstract: Recent high-performing text embedding models often output high-dimensional real-valued vectors, resulting in substantial storage and computational costs

When Knowledge Is Not Free: Cost-Aware Evidence Selection in Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2606.02245v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) typically assumes that external knowledge is free, but many high-quality sources are paywalled, licensed, restricte

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models

SafetyDGX agent

arXiv:2606.01671v1 Announce Type: new Abstract: In the contemporary epoch of multilingual education, learning idioms provides a fascinating gateway towards creativity, cultural values, historical cont

When Rating Scales Fall Short: LLM-Assisted Discovery of ADHD Signals in Turkish Teacher Narratives

ResearchDGX agent

arXiv:2606.02509v1 Announce Type: new Abstract: Attention Deficit Hyperactivity Disorder (ADHD) is one of the most common neurodevelopmental disorders in childhood, and its diagnosis relies on assessm

Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs

ApplicationsDGX agent

arXiv:2606.00333v1 Announce Type: new Abstract: LLMs increasingly answer questions about taxes, labor protections, healthcare, education, pensions, and administrative procedures, where usefulness ofte

Why Do Self-Harm Prediction Models Struggle to Generalise? Lexical and Semantic Variations in Emergency Department Triage Notes

ResearchDGX agent

arXiv:2606.01678v1 Announce Type: new Abstract: Self-harm presentations to emergency departments (EDs) are strongly associated with higher suicide risk. NLP models have shown robust performance in det

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

Model ReleasesDGX agent

arXiv:2606.01276v1 Announce Type: new Abstract: Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs)

1 Jun 2026

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

ResearchDGX agent

arXiv:2603.07751v2 Announce Type: replace-cross Abstract: Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks

A Padding Method for Enhanced Encoding of Inorganic Structures with Varying Chemical Compositions

ResearchDGX agent

arXiv:2605.30743v1 Announce Type: cross Abstract: Designing novel inorganic materials through generative models remains an important challenge for material science, driven by the complexity and divers

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

Model ReleasesDGX agent

arXiv:2605.31351v1 Announce Type: new Abstract: AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer

AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering

ResearchDGX agent

arXiv:2605.31062v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable performance in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, this app

AI for Monitoring and Classifying Data Used in Research Literature

ResearchDGX agent

arXiv:2605.30582v1 Announce Type: new Abstract: While platforms like Google Scholar and Semantic Scholar track citations for academic papers, no comparable infrastructure exists for monitoring dataset

AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis

Model ReleasesDGX agent

arXiv:2605.30599v1 Announce Type: cross Abstract: Medical knowledge is continuously evolving. This creates a need to update or selectively forget information encoded in already-trained medical LLMs. M

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

Model ReleasesDGX agent

arXiv:2605.30804v1 Announce Type: new Abstract: We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for Engl

Are Full Rollouts Necessary for On-Policy Distillation?

SafetyDGX agent

arXiv:2605.31490v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense teacher feedback along rollouts generated by the student and has emerged as a promising post-training paradi

Are we chasing ghosts? Quantifying unattributable polarization, and attributing the rest to annotator groups

ResearchDGX agent

arXiv:2602.06055v2 Announce Type: replace Abstract: Standard agreement metrics often fail to capture systematic differences in opinion between minority and majority-group annotators, jeopardizing task

Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR

TutorialsDGX agent

arXiv:2605.30912v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers.

Auditing LLM Benchmarks with Item Response Theory

Model ReleasesDGX agent

arXiv:2605.30504v1 Announce Type: new Abstract: LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-base

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Model ReleasesDGX agent

arXiv:2605.31483v1 Announce Type: new Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LL

Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed Tokenization

TutorialsDGX agent

arXiv:2510.20853v2 Announce Type: replace-cross Abstract: Electrophysiological (ExG) signals offer valuable insights into human physiology, yet building foundation models that generalize across everyd

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

Model ReleasesDGX agent

arXiv:2605.31086v1 Announce Type: new Abstract: In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the under

Bounded Behavioral Indistinguishability for Black-Box LLM Distillation

Model ReleasesDGX agent

arXiv:2605.30448v1 Announce Type: cross Abstract: Black-box LLM distillation is usually evaluated as an output-matching problem: a student is considered successful when its responses are semantically

Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing

ApplicationsDGX agent

arXiv:2605.31338v1 Announce Type: new Abstract: Statutory references are central to legal language understanding, but are difficult to process automatically, as they appear in compact and variable sur

Can LLM Teams Play What? Where? When?

Model ReleasesDGX agent

arXiv:2605.30459v1 Announce Type: new Abstract: Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigat

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

Model ReleasesDGX agent

arXiv:2605.30497v1 Announce Type: new Abstract: RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmar

Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement

ApplicationsDGX agent

arXiv:2605.30981v1 Announce Type: new Abstract: Autoregressive language models frequently degrade during long-horizon generation, producing repetitive text, losing instruction adherence, and exhibitin

← Previous
1…5051525354…129
Next →