AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Research

Vector RAG vs LLM-Compiled Wiki: A Preregistered Comparison on a Small Multi-Domain Research

DGX agent

arXiv:2605.18490v1 Announce Type: new Abstract: We preregistered a comparison of two ways to help an LLM answer questions over a small research corpus: a single-round Vector RAG system and an LLM-comp

researcharxiv-cs-cl
19 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

DGX agent

arXiv:2605.17467v1 Announce Type: new Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliabil

model-releasesarxiv-cs-cl
19 May 2026
Research

Vidya: An AI-Driven Modular Pipeline for Archival Automation and Semantic Metadata Enrichment

DGX agent

arXiv:2605.16338v1 Announce Type: cross Abstract: The large-scale digitization of historical archives has created a paradox: 'dark data'-digital objects lacking metadata for retrieval. Manual archival

researcharxiv-cs-cl
19 May 2026
Model Releases

We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong

DGX agent

arXiv:2509.22510v3 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) is the ability to satisfy desired objectives during generation, which is critical for trustworthy deployme

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

DGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings

DGX agent

arXiv:2605.16288v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in clinical and care settings. This exploratory study investigates whether LLMs exhibit sycophantic

model-releasesarxiv-cs-cl
19 May 2026
Tutorials

When TableQA Meets Noise: A Dual Denoising Framework for Complex Questions and Large-scale Tables

DGX agent

arXiv:2509.17680v2 Announce Type: replace Abstract: Table question answering (TableQA) is a fundamental task in natural language processing (NLP). The strong reasoning capabilities of large language m

tutorialsarxiv-cs-cl
19 May 2026
Safety

White-Box Sensitivity Auditing with Steering Vectors

DGX agent

arXiv:2601.16398v2 Announce Type: replace-cross Abstract: Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of

safetyarxiv-cs-cl
19 May 2026
Research

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

DGX agent

arXiv:2511.06516v3 Announce Type: replace Abstract: Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering

researcharxiv-cs-cl
19 May 2026
Applications

Adesua: Development and Feasibility Study of an AI WhatsApp Bot for Science Learning in West Africa

DGX agent

arXiv:2605.15376v1 Announce Type: new Abstract: Sub-Saharan Africa faces persistently high student-teacher ratios and shortages of qualified teachers, limiting students' access to personalized learnin

applicationsarxiv-cs-cl
18 May 2026
Model Releases

AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions

DGX agent

arXiv:2601.03707v2 Announce Type: replace Abstract: Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and suffici

model-releasesarxiv-cs-cl
18 May 2026
Research

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

DGX agent

arXiv:2605.15864v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) often produce self-reflective statements like 'let me check the figure again' during reasoning. Do such statements trigg

researcharxiv-cs-cl
18 May 2026
Research

Artificial Aphasias in Lesioned Language Models

DGX agent

arXiv:2605.16222v1 Announce Type: new Abstract: Aphasias, selective language impairments which can arise from brain damage, reveal the functional organization of human language by providing causal lin

researcharxiv-cs-cl
18 May 2026
Model Releases

Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering

DGX agent

arXiv:2605.15362v1 Announce Type: new Abstract: Half a billion citation edges extracted from 100.7 million Ukrainian court decisions reveal that judicial citation structure encodes legal domain bounda

model-releasesarxiv-cs-cl
18 May 2026
Research

Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities

DGX agent

arXiv:2601.21702v3 Announce Type: replace-cross Abstract: We consider Representation Misdirection (RM), a class of large language model (LLM) unlearning methods that achieve forgetting by redirecting

researcharxiv-cs-cl
18 May 2026
Agents

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge

DGX agent

arXiv:2605.15815v1 Announce Type: cross Abstract: Code agents increasingly help developers work with unfamiliar repositories, but every such task depends on a costly prerequisite: bootstrapping the re

agentsarxiv-cs-cl
18 May 2026
Applications

Calibrating LLMs with Semantic-level Reward

DGX agent

arXiv:2605.15588v1 Announce Type: new Abstract: As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate w

applicationsarxiv-cs-cl
18 May 2026
Model Releases

Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction

DGX agent

arXiv:2605.16077v1 Announce Type: new Abstract: Accurate assessment of cognitive decline from spontaneous speech remains challenging due to limited dataset size and class imbalance. In this work, we p

model-releasesarxiv-cs-cl
18 May 2026
Research

Capability Conditioned Scaffolding for Professional Human LLM Collaboration

DGX agent

arXiv:2605.15404v1 Announce Type: new Abstract: Large language model personalization typically adapts outputs to user preferences and style but does not account for differences in user evaluation capa

researcharxiv-cs-cl
18 May 2026
Tutorials

Contexting as Recommendation: Evolutionary Collaborative Filtering for Context Engineering

DGX agent

arXiv:2605.15721v1 Announce Type: new Abstract: Large Language Models (LLMs) are highly sensitive to their input contexts, motivating the development of automated context engineering. However, existin

tutorialsarxiv-cs-cl
18 May 2026
Research

Conversations in Space: Structuring Non-Linear LLM Interactions on a Canvas

DGX agent

arXiv:2605.15848v1 Announce Type: cross Abstract: Conversational interfaces powered by large language models (LLMs) are widely used for ideation and analysis, yet their linear structure limits explora

researcharxiv-cs-cl
18 May 2026
Model Releases

CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency

DGX agent

arXiv:2512.00417v5 Announce Type: replace Abstract: This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large

model-releasesarxiv-cs-cl
18 May 2026
Applications

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory

DGX agent

arXiv:2605.15990v1 Announce Type: new Abstract: Tremendous efforts have been put into evaluating the inclusivity and effectiveness of AI systems across cultures. However, the cultural capabilities con

applicationsarxiv-cs-cl
18 May 2026
Model Releases

DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection

DGX agent

arXiv:2605.15518v1 Announce Type: new Abstract: The effective detection and governance of Large Language Model (LLM) generated content has become increasingly critical due to the growing risk of misus

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory

DGX agent

arXiv:2605.15759v1 Announce Type: new Abstract: Large language model (LLM) agents require long-term memory to leverage information from past interactions. However, existing memory systems often face a

model-releasesarxiv-cs-cl
18 May 2026
Research

DiscoExplorer: An Open Interface for the Study of Multilingual Discourse Relations

DGX agent

arXiv:2605.15304v1 Announce Type: new Abstract: The relations connecting propositions in discourse such as cause (A because B) or concession (A although B) are a subject of intense interest in Computa

researcharxiv-cs-cl
18 May 2026
Tutorials

DiscussLLM: Teaching Large Language Models When to Speak

DGX agent

arXiv:2508.18167v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generating human-like text, yet they largely operate as

tutorialsarxiv-cs-cl
18 May 2026
Model Releases

Dynamic Chunking for Diffusion Language Models

DGX agent

arXiv:2605.15676v1 Announce Type: new Abstract: Block discrete diffusion language models factorize a sequence autoregressively over fixed-size positional blocks, decoupling within-block parallel denoi

model-releasesarxiv-cs-cl
18 May 2026
Safety

Eskwai for Students: Generative AI Assistant for Legal Education in Ghana

DGX agent

arXiv:2605.15380v1 Announce Type: new Abstract: Recent advances in generative AI have shown their potential to be leveraged for legal education. Yet, work on the development and deployment of such sys

safetyarxiv-cs-cl
18 May 2026
Model Releases

Evaluating Chinese Ambiguity Understanding in Large Language Models

DGX agent

arXiv:2605.15635v1 Announce Type: new Abstract: Linguistic ambiguity is critical to the robustness of Large Language Models (LLMs), yet existing research focuses mostly on English, with limited attent

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

Few-Shot Large Language Models for Actionable Triage Categorization of Online Patient Inquiries

DGX agent

arXiv:2605.15680v1 Announce Type: new Abstract: Online patient inquiries are often informal, incomplete, and written before professional assessment, yet they must still be routed to an appropriate lev

model-releasesarxiv-cs-cl
18 May 2026
Research

Few-Step Diffusion Language Models via Trajectory Self-Distillation

DGX agent

arXiv:2602.12262v3 Announce Type: replace Abstract: Diffusion large language models (DLLMs) have emerged as powerful generative models with the promise of fast text generation through parallel decodin

researcharxiv-cs-cl
18 May 2026
Model Releases

FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models

DGX agent

arXiv:2605.15482v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being applied to financial analysis, reporting, investment decision support, risk management, compliance,

model-releasesarxiv-cs-cl
18 May 2026
Safety

FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures

DGX agent

arXiv:2604.05966v2 Announce Type: replace Abstract: Financial reporting systems increasingly leverage Large Language Models (LLMs) to extract and summarize corporate disclosures. However, most existin

safetyarxiv-cs-cl
18 May 2026
Research

Fluency and Faithfulness in Human and Machine Literary Translation

DGX agent

arXiv:2605.15282v1 Announce Type: new Abstract: Literary translation requires balancing target-language fluency with faithfulness to the source. Recent large language models (LLMs) often produce fluen

researcharxiv-cs-cl
18 May 2026
Model Releases

ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation

DGX agent

arXiv:2605.15794v1 Announce Type: new Abstract: We present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

GiLT: Augmenting Transformer Language Models with Dependency Graphs

DGX agent

arXiv:2605.15562v1 Announce Type: new Abstract: Augmenting Transformers with linguistic structures effectively enhances the syntactic generalization performance of language models. Previous work in th

model-releasesarxiv-cs-cl
18 May 2026
Applications

Greedy or not, here I come: Language production under vocabulary constraints in humans and resource-rational models

DGX agent

arXiv:2605.15365v1 Announce Type: new Abstract: Communicating using only a limited vocabulary is a common but challenging cognitive phenomenon, requiring an ideal communicator to plan carefully to opt

applicationsarxiv-cs-cl
18 May 2026
Research

Hallucinations are inevitable but can be made statistically negligible

DGX agent

arXiv:2502.12187v3 Announce Type: replace Abstract: Hallucinations, a phenomenon where a language model (LM) generates nonfactual content, pose a significant challenge to the practical deployment of L

researcharxiv-cs-cl
18 May 2026
Research

Improving Cross-Cultural Survey Simulation with Calibrated Value Personas

DGX agent

arXiv:2605.16193v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate human opinions and survey responses, but their ability to reproduce population responses

researcharxiv-cs-cl
18 May 2026
Model Releases

Introducing MELI: the Mandarin-English Language Interview Corpus

DGX agent

arXiv:2603.27043v2 Announce Type: replace Abstract: We introduce the Mandarin-English Language Interview (MELI) Corpus, an open-source resource of 29.8 hours of speech from 51 Mandarin-English bilingu

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

Judge Circuits

DGX agent

arXiv:2605.16023v1 Announce Type: new Abstract: LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its

model-releasesarxiv-cs-cl
18 May 2026
Safety

Linked Multi-Model Data on Russian Domestic and Foreign Policy Speeches

DGX agent

arXiv:2605.15886v1 Announce Type: new Abstract: This paper introduces a dataset of interlinked multimodal political communications from the Russian government, addressing persistent deficiencies in th

safetyarxiv-cs-cl
18 May 2026
Model Releases

Measuring Maximum Activations in Open Large Language Models

DGX agent

arXiv:2605.15572v1 Announce Type: new Abstract: The dynamic range of activations is a first-order constraint for low-bit quantization, activation scaling, and stable LLM inference. Prior work characte

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

MHGraphBench: Knowledge Graph-Grounded Benchmarking of Mental Health Knowledge in Large Language Models

DGX agent

arXiv:2605.15589v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in the mental health domain, yet it remains unclear how well they capture related biomedical knowledg

model-releasesarxiv-cs-cl
18 May 2026
Research

Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language

DGX agent

arXiv:2511.15887v2 Announce Type: replace Abstract: Our ability to interpret others' mental states through nonverbal cues (NVCs) is fundamental to our survival and social cohesion. While existing Theo

researcharxiv-cs-cl
18 May 2026
Local Ai

Multi-Level Contextual Token Relation Modeling for Machine-Generated Text Detection

DGX agent

arXiv:2605.16107v1 Announce Type: new Abstract: Machine-generated texts (MGTs) pose risks such as disinformation and phishing, underscoring the need for reliable detection. Metric-based methods, which

local-aiarxiv-cs-cl
18 May 2026
Research

Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance

DGX agent

arXiv:2605.15436v1 Announce Type: new Abstract: This paper presents a comprehensive analysis of neural activation patterns across six distinct large language model (LLM) architectures, examining their

researcharxiv-cs-cl
18 May 2026
← Previous
1…9192939495…162
Next →