AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Safety

Safe Evolution with Circuit Anchors

DGX agent

arXiv:2608.05158v1 Announce Type: new Abstract: In biological evolution, unconstrained mutation can lead to catastrophic outcomes: organisms may evolve enhanced capabilities while losing essential fun

safetyarxiv-cs-cl
7 Aug 2026
Model Releases

Scaffold-Mediated Post-Training: Co-Evolving Model Parameters and Procedural Scaffold Graphs

DGX agent
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

arXiv:2608.05156v1 Announce Type: new Abstract: Post-training of large language models optimizes only parameters, while inference-time procedural scaffolds are typically designed independently of para

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters

DGX agent

arXiv:2608.05161v1 Announce Type: new Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining rema

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

Shrinking the Generation-Verification Gap with Weak Verifiers

DGX agent

arXiv:2506.18203v3 Announce Type: replace Abstract: Verifiers can improve language model capabilities by scoring and ranking responses from generated candidates. Currently, high-quality verifiers are

model-releasesarxiv-cs-cl
7 Aug 2026
Research

Sparse Mutual Information Graph Averaging for Improving Random Indexing Embeddings

DGX agent

arXiv:2608.05724v1 Announce Type: new Abstract: Sparse word embedding pipelines can avoid dense co-occurrence matrix materialization, dense factorization, and gradient training while still relying on

researcharxiv-cs-cl
7 Aug 2026
Tutorials

STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughts

DGX agent

arXiv:2602.14265v3 Announce Type: replace Abstract: Inference-Time-Compute (ITC) methods like Best-of-n and Tree-of-Thoughts are meant to produce output candidates that are both high-quality and diver

tutorialsarxiv-cs-cl
7 Aug 2026
Safety

Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges

DGX agent

arXiv:2405.15604v4 Announce Type: replace Abstract: Text generation has become more accessible than ever, and the growing interest in these systems, especially those using large language models, has s

safetyarxiv-cs-cl
7 Aug 2026
Model Releases

The Bitter Lesson of Tool Calling

DGX agent

arXiv:2608.06370v1 Announce Type: new Abstract: Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by

model-releasesarxiv-cs-cl
7 Aug 2026
Research

The interface of intonation and lexical tone: Boundary phenomena in Mandarin varieties

DGX agent

arXiv:2608.05364v1 Announce Type: new Abstract: This chapter explores the intricate interplay between intonation and tone in Mandarin Chinese varieties, focusing on f0, the primary acoustic cue for bo

researcharxiv-cs-cl
7 Aug 2026
Agents

The Vulnerability With No CVE: Managing Persistent Gaps Between Mandate and Authority in AI Coding Agents

DGX agent

arXiv:2608.05884v1 Announce Type: cross Abstract: Existing guidance identifies excessive agency, excessive permission, weak task-bound authorization, and inadequate agent controls as important risks.

agentsarxiv-cs-cl
7 Aug 2026
Research

Training-Free Token-Level Steering for LLM Personalized Co-Writing

DGX agent

arXiv:2608.06069v1 Announce Type: new Abstract: While Large Language Models (LLMs) show great promise for personalization, they often lack specialized domain knowledge. Conventional solutions like fin

researcharxiv-cs-cl
7 Aug 2026
Research

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

DGX agent

arXiv:2608.05576v1 Announce Type: new Abstract: When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizab

researcharxiv-cs-cl
7 Aug 2026
Model Releases

Where Privacy Risk Lives in English-Source Multilingual RAG: A Stage-Decomposed Audit Across Five Query Languages

DGX agent

arXiv:2608.05163v1 Announce Type: new Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information. We test th

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

Who Checks the Citations? Benchmarking Legal Hallucination Detection

DGX agent

arXiv:2606.21155v2 Announce Type: replace Abstract: Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictio

model-releasesarxiv-cs-cl
7 Aug 2026
Model Releases

Zero-Shot Multi-Disease Labeling of Chest, Abdomen, and Pelvis CT Reports Using Open-Weight Large Language Models: The Effect of Labeling Conventions

DGX agent

arXiv:2506.03259v3 Announce Type: replace Abstract: Purpose: To compare five lightweight open-weight large language models (LLMs) with a rule-based algorithm (RBA) and fine-tuned RadBERT for zero-shot

model-releasesarxiv-cs-cl
7 Aug 2026
Research

A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy

DGX agent

arXiv:2608.04808v1 Announce Type: new Abstract: Part-of-speech tagging for low-resource languages remains challenging due to limited annotated data, especially for linguistically complex languages. Ga

researcharxiv-cs-cl
6 Aug 2026
Model Releases

Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems

DGX agent

arXiv:2608.04746v1 Announce Type: new Abstract: LLM agents that persist across sessions accumulate stored memories whose validity varies enormously by content type, yet existing memory architectures t

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study

DGX agent

arXiv:2608.04576v1 Announce Type: new Abstract: Humanitarian reports are long, noisy, and multi-topic, making it difficult to consolidate decision-relevant causal evidence. We present a ReliefWeb stud

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

DataRx: Missingness-Aware Sampling for Safer Large Language Model Task-Specific Fine-Tuning

DGX agent

arXiv:2608.04322v1 Announce Type: new Abstract: Task-specific fine-tuning can improve the performance of large language models (LLMs) on downstream tasks. However, our study reveals that task-specific

safetyarxiv-cs-cl
6 Aug 2026
Research

DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models

DGX agent

arXiv:2608.04477v1 Announce Type: cross Abstract: Cloud-based language model services routinely process prompts containing sensitive information. Obfuscation-based defenses---including ObfusLM, Sentin

researcharxiv-cs-cl
6 Aug 2026
Model Releases

DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots

DGX agent

arXiv:2608.05004v1 Announce Type: new Abstract: Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including 'delusi

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Do Language Models Know Their Slang? Queer Slang Understanding in User-Generated Content

DGX agent

arXiv:2608.04847v1 Announce Type: new Abstract: Despite its cultural relevance and diffusion, queer slang remains underrepresented in Natural Language Processing research. Towards addressing this gap,

researcharxiv-cs-cl
6 Aug 2026
Research

Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations

DGX agent

arXiv:2604.00209v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they frequently violate contextual privacy by disclosing private

researcharxiv-cs-cl
6 Aug 2026
Research

Document Optimization for Black-Box Retrieval via Reinforcement Learning

DGX agent

arXiv:2604.05087v3 Announce Type: replace Abstract: Document expansion is a classical technique for improving retrieval quality, and is attractive since it shifts computation offline, avoiding additio

researcharxiv-cs-cl
6 Aug 2026
Research

Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?

DGX agent

arXiv:2608.04928v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning offers a window into the decision-making of large language models (LLMs), which can be monitored for target behaviors b

researcharxiv-cs-cl
6 Aug 2026
Research

EdgeLM: Edge Demonstrations for Language Models' Table Understanding

DGX agent

arXiv:2608.04390v1 Announce Type: new Abstract: Large language models (LLMs) perform table-centric prediction through in-context learning, making demonstration selection critical to performance. Exist

researcharxiv-cs-cl
6 Aug 2026
Model Releases

Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks

DGX agent

arXiv:2608.04286v1 Announce Type: new Abstract: Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations

model-releasesarxiv-cs-cl
6 Aug 2026
Agents

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot

DGX agent

arXiv:2608.04709v1 Announce Type: new Abstract: This paper presents EmpaAva, to our knowledge the first open-source, agentic 3D-avatar empathetic chatbot, which carries empathetic response generation

agentsarxiv-cs-cl
6 Aug 2026
Model Releases

Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs

DGX agent

arXiv:2608.04488v1 Announce Type: new Abstract: Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Enhancing Trustworthy Clinical Diagnosis Decision-Making in Large Language Models via Etiology-Aware Attention Supervision

DGX agent

arXiv:2508.00285v2 Announce Type: replace Abstract: Objective: Large Language Models (LLMs) have demonstrated strong capabilities in medical text understanding and generation. However, their trustwort

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

DGX agent

arXiv:2608.04885v1 Announce Type: cross Abstract: Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validat

safetyarxiv-cs-cl
6 Aug 2026
Research

Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

DGX agent

arXiv:2608.04646v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the

researcharxiv-cs-cl
6 Aug 2026
Research

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

DGX agent

arXiv:2608.04899v1 Announce Type: new Abstract: Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as

researcharxiv-cs-cl
6 Aug 2026
Model Releases

GEB-Bench: Abstract Structures Told in Many Voices

DGX agent

arXiv:2608.04111v1 Announce Type: cross Abstract: Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstr

model-releasesarxiv-cs-cl
6 Aug 2026
Research

German parties shifted towards intuition-based rhetoric after the far right's parliamentary breakthrough

DGX agent

arXiv:2608.05075v1 Announce Type: new Abstract: The spread of misinformation is widely perceived as a threat to democratic deliberation, yet how political elites' rhetorical commitments to truth shift

researcharxiv-cs-cl
6 Aug 2026
Model Releases

IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

DGX agent

arXiv:2608.04703v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where an

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

K-EXAONE 2.0 Technical Report

DGX agent

arXiv:2608.04505v1 Announce Type: new Abstract: This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward glo

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention

DGX agent

arXiv:2608.04678v1 Announce Type: new Abstract: Papers 1-2 of the Kathleen series showed that a byte-level, attention-free architecture built from a wavetable encoder and multi-scale reverberant state

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

Language Models Generalize to Human-like Word Order Preferences

DGX agent

arXiv:2608.05028v1 Announce Type: new Abstract: A central question in language acquisition is whether linguistic biases can emerge from general learning mechanisms operating over underdetermined input

safetyarxiv-cs-cl
6 Aug 2026
Model Releases

Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language

DGX agent

arXiv:2608.04186v1 Announce Type: new Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs).

model-releasesarxiv-cs-cl
6 Aug 2026
Agents

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

DGX agent

arXiv:2608.04056v1 Announce Type: new Abstract: When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP sys

agentsarxiv-cs-cl
6 Aug 2026
Research

Leveraging Machine Learning to Gain Insights on Quantum Thermodynamic Entropy

DGX agent

arXiv:2305.06177v1 Announce Type: cross Abstract: We present a thermodynamic analysis of a quantum engine that uses a single quantum particle as its working fluid, inspired by Szilard's classical sing

researcharxiv-cs-cl
6 Aug 2026
Research

LLM-based Vulnerability Discovery in Business Process Documentation

DGX agent

arXiv:2608.04271v1 Announce Type: cross Abstract: Just like software and hardware, business processes are susceptible to vulnerabilities that can lead to product quality issues, delays, and increased

researcharxiv-cs-cl
6 Aug 2026
Safety

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

DGX agent

arXiv:2606.18709v2 Announce Type: replace Abstract: Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item mea

safetyarxiv-cs-cl
6 Aug 2026
Model Releases

MathDebugger: Detecting and Diagnosing Errors in Synthetic Mathematical Data

DGX agent

arXiv:2502.19058v2 Announce Type: replace Abstract: Synthetic mathematical data has become an important resource for scaling the reasoning capabilities of large language models, yet errors in generate

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap

DGX agent

arXiv:2608.04160v1 Announce Type: new Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so th

model-releasesarxiv-cs-cl
6 Aug 2026
Research

Neighborhood-Aware Dual Biomedical Entity Linking

DGX agent

arXiv:2608.04144v1 Announce Type: cross Abstract: Biomedical entity linking grounds mentions in clinical and scientific text to entities in a curated knowledge base (KB) with ontological structure, wh

researcharxiv-cs-cl
6 Aug 2026
Applications

Neurocomputational Mechanisms of Syntactic Transfer in Bilingual Sentence Production

DGX agent

arXiv:2601.18056v2 Announce Type: replace Abstract: We discuss the benefits of incorporating oscillatory neural mechanisms into the study of bilingual production errors and their traditionally documen

applicationsarxiv-cs-cl
6 Aug 2026
← Previous
1…56789…160
Next →