AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation

DGX agent

arXiv:2607.20862v1 Announce Type: new Abstract: At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative

safetyarxiv-cs-cl
24 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages

DGX agent

arXiv:2607.21016v1 Announce Type: new Abstract: Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping aw

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

DGX agent

arXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

DGX agent

arXiv:2603.11838v2 Announce Type: replace Abstract: Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome dur

model-releasesarxiv-cs-cl
24 Jul 2026
Research

Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content

DGX agent

arXiv:2607.20446v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used by students to generate natural language responses and program code, there is growing interest in

researcharxiv-cs-cl
24 Jul 2026
Model Releases

Domyn-Small: A European 10B Reasoning Language Model

DGX agent

arXiv:2607.20448v1 Announce Type: new Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an i

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

DGX agent

arXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1

DGX agent

arXiv:2607.20589v1 Announce Type: new Abstract: Persona simulation involves utilizing large language models (LLMs) to anticipate human choices or interactions based on specific characteristic informat

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

From a Word-Level Dictionary to Sentence-Level Semantics: Multilingual Grievance Labelling with Contextual Models

DGX agent

arXiv:2607.20946v1 Announce Type: new Abstract: Grievance is one of the warning signs analysts look for when assessing threats of violence. It is increasingly measured at scale from online text, most

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries

DGX agent

arXiv:2607.20757v1 Announce Type: cross Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus

DGX agent

arXiv:2607.20443v1 Announce Type: new Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipel

model-releasesarxiv-cs-cl
24 Jul 2026
Tutorials

Gumbel Distillation for Parallel Text Generation

DGX agent

arXiv:2603.22216v2 Announce Type: replace Abstract: The slow, sequential nature of autoregressive (AR) language models has driven the adoption of parallel decoding methods. However, these non-AR model

tutorialsarxiv-cs-cl
24 Jul 2026
Agents

Human-in-the-Loop Large Language Model Framework for Identification of Cutaneous Immune-Related Adverse Events

DGX agent

arXiv:2607.20428v1 Announce Type: new Abstract: This study evaluated a retrieval-augmented, multi-agent large language model (LLM)-driven, human-in-the-loop framework for detecting cutaneous immune-re

agentsarxiv-cs-cl
24 Jul 2026
Applications

Learning to Detect UI Principle Violations via Reinforcement Learning

DGX agent

arXiv:2607.20690v1 Announce Type: new Abstract: Small language models and coding agents increasingly generate web front-end code, yet their outputs are typically evaluated primarily for functional cor

applicationsarxiv-cs-cl
24 Jul 2026
Model Releases

LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

DGX agent

arXiv:2607.20872v1 Announce Type: new Abstract: Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, b

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

DGX agent

arXiv:2607.21570v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

MemTools: A Unified Research Framework for Interoperable Agent Memory

DGX agent

arXiv:2607.21404v1 Announce Type: new Abstract: While memory systems are essential for agent architectures, pervasive architectural fragmentation restricts systematic research. Existing implementation

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models

DGX agent

arXiv:2607.20442v1 Announce Type: new Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in Jul

model-releasesarxiv-cs-cl
24 Jul 2026
Research

news-crawler-LM: A Small Long-Context Model For High-Quality News Crawling

DGX agent

arXiv:2607.21284v1 Announce Type: new Abstract: Extracting structured content from news pages remains challenging due to heterogeneous HTML layouts, inconsistent markup, and substantial boilerplate su

researcharxiv-cs-cl
24 Jul 2026
Research

Non-Zipfian Distribution of Stopwords or Function Words and Subset Selection Models

DGX agent

arXiv:2603.04691v2 Announce Type: replace Abstract: Stopwords and function words are relatively less informative for the content of a language and more often play a structural role in a sentence. Stop

researcharxiv-cs-cl
24 Jul 2026
Model Releases

Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks

DGX agent

arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single ans

model-releasesarxiv-cs-cl
24 Jul 2026
Applications

Position: Natural Language Should Not Fully Replace Formal Languages

DGX agent

arXiv:2607.20432v1 Announce Type: new Abstract: Recent advances in large language models and their widespread adoption have prompted claims that natural language could entirely replace formal language

applicationsarxiv-cs-cl
24 Jul 2026
Tutorials

PrefReward: Learning User Preference Matrix for Personalized Text Generation

DGX agent

arXiv:2607.21067v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable ability in generating personalized content by leveraging user histories and contextual cues. H

tutorialsarxiv-cs-cl
24 Jul 2026
Research

Progressive Cramming: Reliable Token Compression and What It Reveals

DGX agent

arXiv:2607.21231v1 Announce Type: new Abstract: Token cramming compresses sequences into learned embeddings with near-perfect reconstruction, but fixed token budgets and 99% accuracy thresholds leave

researcharxiv-cs-cl
24 Jul 2026
Model Releases

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

DGX agent

arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is as

model-releasesarxiv-cs-cl
24 Jul 2026
Local Ai

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

DGX agent

arXiv:2607.20833v1 Announce Type: new Abstract: Large language models increasingly rely on long-form reasoning for complex tasks, yet their reasoning traces may drift away from the supplied context wh

local-aiarxiv-cs-cl
24 Jul 2026
Model Releases

REGARD: Regional Affective Differences in Large Language Models

DGX agent

arXiv:2607.20722v1 Announce Type: new Abstract: Large language models trained and aligned within different linguistic and regional ecosystems may frame the same political, cultural, and geopolitical e

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Rushes: A Human Preference Dataset for Pluralistic Alignment

DGX agent

arXiv:2607.20767v1 Announce Type: new Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collect

model-releasesarxiv-cs-cl
24 Jul 2026
Agents

Sample-Efficient Learning from Agent Experience

DGX agent

arXiv:2607.21051v1 Announce Type: new Abstract: Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedbac

agentsarxiv-cs-cl
24 Jul 2026
Tutorials

SCoPE: Shift-Aware Speaker-Conditioned Priors for Emotion Recognition in Conversations

DGX agent

arXiv:2607.20445v1 Announce Type: new Abstract: In conversations, human emotions are transient; however, they tend to persist across multiple utterances. For example, we rarely switch instantly betwee

tutorialsarxiv-cs-cl
24 Jul 2026
Research

Semantic Field Theory: Historical Origin, Higher-Order Interaction, and Stabilized Semantic Inference

DGX agent

arXiv:2607.20451v1 Announce Type: new Abstract: Semantic Field Theory (SFT) has developed from a philosophical critique of strong anti-formalist readings of language games into a proposed computationa

researcharxiv-cs-cl
24 Jul 2026
Research

ShriNep@EEUCA 2026: RAKSHAK - Multi-Task DeBERTa with Rationale Distillation and Jigsaw-Augmented Training for Toxic Intent Classification

DGX agent

arXiv:2607.20450v1 Announce Type: new Abstract: This paper presents two systems for the GameTox Shared Task at the Workshop on EEUCA at ACL 2026, which requires classifying World of Tanks chat utteran

researcharxiv-cs-cl
24 Jul 2026
Model Releases

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

DGX agent

arXiv:2607.09999v2 Announce Type: replace Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-catego

model-releasesarxiv-cs-cl
24 Jul 2026
Agents

Skill-Contracted Agents for Evidence-Aware Materials Literature Analysis

DGX agent

arXiv:2607.20431v1 Announce Type: new Abstract: Materials science literature analysis requires simultaneous attention to composition, processing, characterization, and property relationships, yet conv

agentsarxiv-cs-cl
24 Jul 2026
Model Releases

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

DGX agent

arXiv:2607.15557v4 Announce Type: replace Abstract: Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Pub

model-releasesarxiv-cs-cl
24 Jul 2026
Research

Surprisal Theory is Tautological (without Rational Grounding)

DGX agent

arXiv:2607.21574v1 Announce Type: new Abstract: Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language m

researcharxiv-cs-cl
24 Jul 2026
Model Releases

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

DGX agent

arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring pro

model-releasesarxiv-cs-cl
24 Jul 2026
Safety

thaulab@EEUCA 2026: Who Said What to Whom? A Targeting-Aware Neural-Symbolic Pipeline for Gaming Toxicity Detection

DGX agent

arXiv:2607.20447v1 Announce Type: new Abstract: This paper describes our system for the EEUCA 2026 Shared Task on toxicity classification in gaming chat. We implement a three-stage pipeline combining

safetyarxiv-cs-cl
24 Jul 2026
Research

The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning

DGX agent

arXiv:2607.20952v1 Announce Type: cross Abstract: Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assume

researcharxiv-cs-cl
24 Jul 2026
Safety

Token-Level Entropy Reveals Demographic Disparities in Large Language Models

DGX agent

arXiv:2501.19337v5 Announce Type: replace Abstract: A name alone measurably reshapes a language model's next-token distribution before a single token is sampled. We measure full-vocabulary Shannon ent

safetyarxiv-cs-cl
24 Jul 2026
Applications

TopoGuard: Graph Theory Based Defenses Against Split-Knowledge Attacks on RAG

DGX agent

arXiv:2607.20437v1 Announce Type: new Abstract: Production Retrieval Augmented Generation (RAG) systems rely on aggregating multiple external documents to answer complex queries. However, the retrieve

applicationsarxiv-cs-cl
24 Jul 2026
Research

Transformer-Assisted LLM-Based Source Code Summarisation: to Enable More Secure Software Development

DGX agent

arXiv:2607.20933v1 Announce Type: cross Abstract: Neural Source Code Summarisation (NSCS) aims to generate natural language summaries of source code to improve developers' and maintainers' understandi

researcharxiv-cs-cl
24 Jul 2026
Local Ai

VibeVoice-ASR-BitNet Technical Report

DGX agent

arXiv:2607.21075v1 Announce Type: cross Abstract: We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantiza

local-aiarxiv-cs-cl
24 Jul 2026
Model Releases

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

DGX agent

arXiv:2607.20425v1 Announce Type: new Abstract: What makes writing 'good' remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how r

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

DGX agent

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

DGX agent

arXiv:2607.21445v1 Announce Type: new Abstract: Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this pap

model-releasesarxiv-cs-cl
24 Jul 2026
Research

Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept

DGX agent

arXiv:2607.20995v1 Announce Type: new Abstract: Distinguishing animate from inanimate concepts in written language requires more than shallow text processing, as it involves recognizing complex select

researcharxiv-cs-cl
24 Jul 2026
Model Releases

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

DGX agent

arXiv:2607.21535v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models

model-releasesarxiv-cs-cl
24 Jul 2026
← Previous
1…2526272829…161
Next →