AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete

DGX agent

arXiv:2606.00048v1 Announce Type: cross Abstract: Prior research has established that instruction-tuned large language models exhibit left-of-center political bias, measured exclusively through abstra

model-releasesarxiv-cs-cl
2 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

The Social Cost of Intelligence: Emergence, Propagation, and Amplification of Stereotypical Bias in Multi-Agent Systems

DGX agent

arXiv:2510.10943v2 Announce Type: replace-cross Abstract: Bias in large language models (LLMs) remains a persistent challenge, often leading to stereotyping and unfair treatment across social groups.

safetyarxiv-cs-cl
2 Jun 2026
Research

Thinking Economically: A Hierarchical Framework for Adaptive-Complexity Reasoning in LLMs

DGX agent

arXiv:2606.01168v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has significantly enhanced LLM reasoning, yet often incurs substantial computational overhead due to 'overthinking': generating e

researcharxiv-cs-cl
2 Jun 2026
Tutorials

ToMAP: Training Opponent-Aware LLM Persuaders with Theory of Mind

DGX agent

arXiv:2505.22961v3 Announce Type: replace Abstract: Large language models (LLMs) have shown promising potential in persuasion, but existing works on training LLM persuaders are still preliminary. Nota

tutorialsarxiv-cs-cl
2 Jun 2026
Safety

Toward Responsible and Epistemically Grounded Multilingual LLMs for Computational Social Science and Humanities

DGX agent

arXiv:2606.00596v1 Announce Type: new Abstract: Large language models have rapidly evolved in multilingual competence and reasoning capacity, enabling their integration into Social Sciences and Humani

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

DGX agent

arXiv:2606.00919v1 Announce Type: new Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - re

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

DGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

model-releasesarxiv-cs-cl
2 Jun 2026
Research

Training Prompt Matters: State-Adaptive Optimization for Robust Fine-Tuning

DGX agent

arXiv:2606.01967v1 Announce Type: new Abstract: While prompt engineering is instrumental in maximizing the capabilities of Large Language Models (LLMs) during inference, the role of prompts during tra

researcharxiv-cs-cl
2 Jun 2026
Research

Transferable Self-Harm Surveillance from Emergency Department Triage Notes Using an Evidence-Augmented Machine Learning Approach

DGX agent

arXiv:2606.02545v1 Announce Type: new Abstract: Self-harm is a major public health concern, but current surveillance relying on hospital presentations is inadequate due to the low sensitivity of diagn

researcharxiv-cs-cl
2 Jun 2026
Tutorials

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher

DGX agent

arXiv:2606.01000v1 Announce Type: cross Abstract: Weak-to-strong generalization studies how to improve a strong student using supervision from a weaker teacher when reliable labels are scarce. We view

tutorialsarxiv-cs-cl
2 Jun 2026
Safety

Trust Region On-Policy Distillation

DGX agent

arXiv:2606.01249v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

DGX agent

arXiv:2606.01456v1 Announce Type: cross Abstract: Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sale

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation

DGX agent

arXiv:2606.02320v1 Announce Type: new Abstract: Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmar

model-releasesarxiv-cs-cl
2 Jun 2026
Research

Uncovering Temporal Framing in the News

DGX agent

arXiv:2606.00294v1 Announce Type: new Abstract: Temporal language does more than place events on a timeline. In news discourse, references to the past, present, and future can function as rhetorical d

researcharxiv-cs-cl
2 Jun 2026
Model Releases

UniD^3: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning

DGX agent

arXiv:2606.01394v1 Announce Type: new Abstract: Systematic characterization of drug-disease relationships is essential for drug discovery and repurposing, yet is hindered by the heterogeneity and rapi

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

Unified Context Evolution for LLM Agents

DGX agent

arXiv:2606.02304v1 Announce Type: new Abstract: LLM-based agents can solve multi-step interactive tasks by combining reasoning with environment feedback, yet each episode starts from the same fixed co

agentsarxiv-cs-cl
2 Jun 2026
Model Releases

Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention

DGX agent

arXiv:2606.01243v1 Announce Type: new Abstract: Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over ex

model-releasesarxiv-cs-cl
2 Jun 2026
Research

Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning

DGX agent

arXiv:2606.02020v1 Announce Type: new Abstract: This paper investigates the entropy dynamics of Chain-of-Thought (CoT) and uncovers a consistent two-phase structure: an Uncertainty Region of explorati

researcharxiv-cs-cl
2 Jun 2026
Safety

VERA: Variational Inference Framework for Jailbreaking Large Language Models

DGX agent

arXiv:2506.22666v3 Announce Type: replace-cross Abstract: The rise of API-only access to state-of-the-art LLMs highlights the need for effective black-box jailbreak methods to identify model vulnerabi

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

DGX agent

arXiv:2510.22276v3 Announce Type: replace-cross Abstract: Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing Engl

model-releasesarxiv-cs-cl
2 Jun 2026
Research

WAXAL-NET: Finetuned Edge ASR Across 19 African Languages

DGX agent

arXiv:2606.02375v1 Announce Type: new Abstract: We evaluate whether compact domain-specialized ASR models can outperform massively multilingual foundation models for conversational African speech acro

researcharxiv-cs-cl
2 Jun 2026
Model Releases

What to Format and How: A Benchmark and Workflow Approach for Document Formatting

DGX agent

arXiv:2606.01936v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have opened up new possibilities for automated document formatting. However, real-world formatting often

model-releasesarxiv-cs-cl
2 Jun 2026
Research

When Is 0.1% Enough? Analyzing the Combined Effects of Dimensionality Reduction and Quantization on Text Embedding Compression

DGX agent

arXiv:2606.01074v1 Announce Type: new Abstract: Recent high-performing text embedding models often output high-dimensional real-valued vectors, resulting in substantial storage and computational costs

researcharxiv-cs-cl
2 Jun 2026
Agents

When Knowledge Is Not Free: Cost-Aware Evidence Selection in Retrieval-Augmented Generation

DGX agent

arXiv:2606.02245v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) typically assumes that external knowledge is free, but many high-quality sources are paywalled, licensed, restricte

agentsarxiv-cs-cl
2 Jun 2026
Safety

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models

DGX agent

arXiv:2606.01671v1 Announce Type: new Abstract: In the contemporary epoch of multilingual education, learning idioms provides a fascinating gateway towards creativity, cultural values, historical cont

safetyarxiv-cs-cl
2 Jun 2026
Research

When Rating Scales Fall Short: LLM-Assisted Discovery of ADHD Signals in Turkish Teacher Narratives

DGX agent

arXiv:2606.02509v1 Announce Type: new Abstract: Attention Deficit Hyperactivity Disorder (ADHD) is one of the most common neurodevelopmental disorders in childhood, and its diagnosis relies on assessm

researcharxiv-cs-cl
2 Jun 2026
Applications

Which Institutional Frameworks Do Chatbots Assume? Auditing Jurisdictional Defaults in Multilingual LLMs

DGX agent

arXiv:2606.00333v1 Announce Type: new Abstract: LLMs increasingly answer questions about taxes, labor protections, healthcare, education, pensions, and administrative procedures, where usefulness ofte

applicationsarxiv-cs-cl
2 Jun 2026
Research

Why Do Self-Harm Prediction Models Struggle to Generalise? Lexical and Semantic Variations in Emergency Department Triage Notes

DGX agent

arXiv:2606.01678v1 Announce Type: new Abstract: Self-harm presentations to emergency departments (EDs) are strongly associated with higher suicide risk. NLP models have shown robust performance in det

researcharxiv-cs-cl
2 Jun 2026
Model Releases

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

DGX agent

arXiv:2606.01276v1 Announce Type: new Abstract: Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs)

model-releasesarxiv-cs-cl
2 Jun 2026
Research

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

DGX agent

arXiv:2603.07751v2 Announce Type: replace-cross Abstract: Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks

researcharxiv-cs-cl
1 Jun 2026
Research

A Padding Method for Enhanced Encoding of Inorganic Structures with Varying Chemical Compositions

DGX agent

arXiv:2605.30743v1 Announce Type: cross Abstract: Designing novel inorganic materials through generative models remains an important challenge for material science, driven by the complexity and divers

researcharxiv-cs-cl
1 Jun 2026
Model Releases

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

DGX agent

arXiv:2605.31351v1 Announce Type: new Abstract: AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer

model-releasesarxiv-cs-cl
1 Jun 2026
Research

AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering

DGX agent

arXiv:2605.31062v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable performance in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, this app

researcharxiv-cs-cl
1 Jun 2026
Research

AI for Monitoring and Classifying Data Used in Research Literature

DGX agent

arXiv:2605.30582v1 Announce Type: new Abstract: While platforms like Google Scholar and Semantic Scholar track citations for academic papers, no comparable infrastructure exists for monitoring dataset

researcharxiv-cs-cl
1 Jun 2026
Model Releases

AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis

DGX agent

arXiv:2605.30599v1 Announce Type: cross Abstract: Medical knowledge is continuously evolving. This creates a need to update or selectively forget information encoded in already-trained medical LLMs. M

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

DGX agent

arXiv:2605.30804v1 Announce Type: new Abstract: We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for Engl

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

Are Full Rollouts Necessary for On-Policy Distillation?

DGX agent

arXiv:2605.31490v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense teacher feedback along rollouts generated by the student and has emerged as a promising post-training paradi

safetyarxiv-cs-cl
1 Jun 2026
Research

Are we chasing ghosts? Quantifying unattributable polarization, and attributing the rest to annotator groups

DGX agent

arXiv:2602.06055v2 Announce Type: replace Abstract: Standard agreement metrics often fail to capture systematic differences in opinion between minority and majority-group annotators, jeopardizing task

researcharxiv-cs-cl
1 Jun 2026
Tutorials

Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR

DGX agent

arXiv:2605.30912v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers.

tutorialsarxiv-cs-cl
1 Jun 2026
Model Releases

Auditing LLM Benchmarks with Item Response Theory

DGX agent

arXiv:2605.30504v1 Announce Type: new Abstract: LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-base

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

DGX agent

arXiv:2605.31483v1 Announce Type: new Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LL

model-releasesarxiv-cs-cl
1 Jun 2026
Tutorials

Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed Tokenization

DGX agent

arXiv:2510.20853v2 Announce Type: replace-cross Abstract: Electrophysiological (ExG) signals offer valuable insights into human physiology, yet building foundation models that generalize across everyd

tutorialsarxiv-cs-cl
1 Jun 2026
Model Releases

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

DGX agent

arXiv:2605.31086v1 Announce Type: new Abstract: In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the under

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Bounded Behavioral Indistinguishability for Black-Box LLM Distillation

DGX agent

arXiv:2605.30448v1 Announce Type: cross Abstract: Black-box LLM distillation is usually evaluated as an output-matching problem: a student is considered successful when its responses are semantically

model-releasesarxiv-cs-cl
1 Jun 2026
Applications

Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing

DGX agent

arXiv:2605.31338v1 Announce Type: new Abstract: Statutory references are central to legal language understanding, but are difficult to process automatically, as they appear in compact and variable sur

applicationsarxiv-cs-cl
1 Jun 2026
Model Releases

Can LLM Teams Play What? Where? When?

DGX agent

arXiv:2605.30459v1 Announce Type: new Abstract: Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigat

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

DGX agent

arXiv:2605.30497v1 Announce Type: new Abstract: RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmar

model-releasesarxiv-cs-cl
1 Jun 2026
Applications

Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement

DGX agent

arXiv:2605.30981v1 Announce Type: new Abstract: Autoregressive language models frequently degrade during long-horizon generation, producing repetitive text, losing instruction adherence, and exhibitin

applicationsarxiv-cs-cl
1 Jun 2026
← Previous
1…6364656667…161
Next →