AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Research

Are we chasing ghosts? Quantifying unattributable polarization, and attributing the rest to annotator groups

DGX agent

arXiv:2602.06055v2 Announce Type: replace Abstract: Standard agreement metrics often fail to capture systematic differences in opinion between minority and majority-group annotators, jeopardizing task

researcharxiv-cs-cl
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tutorials

Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR

DGX agent

arXiv:2605.30912v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers.

tutorialsarxiv-cs-cl
1 Jun 2026
Model Releases

Auditing LLM Benchmarks with Item Response Theory

DGX agent

arXiv:2605.30504v1 Announce Type: new Abstract: LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-base

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

DGX agent

arXiv:2605.31483v1 Announce Type: new Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LL

model-releasesarxiv-cs-cl
1 Jun 2026
Tutorials

Beyond Hearing: Learning Task-Agnostic ExG Representations from Earphones via Physiology-Informed Tokenization

DGX agent

arXiv:2510.20853v2 Announce Type: replace-cross Abstract: Electrophysiological (ExG) signals offer valuable insights into human physiology, yet building foundation models that generalize across everyd

tutorialsarxiv-cs-cl
1 Jun 2026
Model Releases

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

DGX agent

arXiv:2605.31086v1 Announce Type: new Abstract: In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the under

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Bounded Behavioral Indistinguishability for Black-Box LLM Distillation

DGX agent

arXiv:2605.30448v1 Announce Type: cross Abstract: Black-box LLM distillation is usually evaluated as an output-matching problem: a student is considered successful when its responses are semantically

model-releasesarxiv-cs-cl
1 Jun 2026
Applications

Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing

DGX agent

arXiv:2605.31338v1 Announce Type: new Abstract: Statutory references are central to legal language understanding, but are difficult to process automatically, as they appear in compact and variable sur

applicationsarxiv-cs-cl
1 Jun 2026
Model Releases

Can LLM Teams Play What? Where? When?

DGX agent

arXiv:2605.30459v1 Announce Type: new Abstract: Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigat

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

DGX agent

arXiv:2605.30497v1 Announce Type: new Abstract: RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmar

model-releasesarxiv-cs-cl
1 Jun 2026
Applications

Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement

DGX agent

arXiv:2605.30981v1 Announce Type: new Abstract: Autoregressive language models frequently degrade during long-horizon generation, producing repetitive text, losing instruction adherence, and exhibitin

applicationsarxiv-cs-cl
1 Jun 2026
Research

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination

DGX agent

arXiv:2605.31058v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Langu

researcharxiv-cs-cl
1 Jun 2026
Safety

Configurable Reward Model for Balanced Safety Alignment

DGX agent

arXiv:2605.30487v1 Announce Type: new Abstract: Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned

safetyarxiv-cs-cl
1 Jun 2026
Safety

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

DGX agent

arXiv:2605.31073v1 Announce Type: new Abstract: Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do

safetyarxiv-cs-cl
1 Jun 2026
Research

Consolidating Rewarded Perturbations for LLM Post-Training

DGX agent

arXiv:2605.31494v1 Announce Type: new Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by

researcharxiv-cs-cl
1 Jun 2026
Research

Context-Free Recognition with Transformers

DGX agent

arXiv:2601.01754v3 Announce Type: replace-cross Abstract: Transformers excel empirically on tasks that process well-formed inputs according to some grammar, such as natural language and code. However,

researcharxiv-cs-cl
1 Jun 2026
Agents

Counterfactual Graph for Multi-Agent LLM Calibration

DGX agent

arXiv:2605.30653v1 Announce Type: new Abstract: Multi-agent LLM systems often treat agreement as evidence: when many agents in a panel give the same answer, that answer is assumed to be more reliable.

agentsarxiv-cs-cl
1 Jun 2026
Research

Cross-Lingual Steering for Figurative Language Generation

DGX agent

arXiv:2605.30443v1 Announce Type: new Abstract: Multilingual large language models can generate figurative language, but whether the internal signals driving this behavior are language-specific or reu

researcharxiv-cs-cl
1 Jun 2026
Model Releases

CSULoRA: Closest Safe Update Low-Rank Adaptation

DGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

model-releasesarxiv-cs-cl
1 Jun 2026
Hardware

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch

DGX agent

arXiv:2511.17826v2 Announce Type: replace-cross Abstract: Deterministic inference is increasingly critical for large language model (LLM) applications such as LLM-as-a-judge evaluation, multi-agent sy

hardwarearxiv-cs-cl
1 Jun 2026
Tutorials

Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection

DGX agent

arXiv:2605.31563v1 Announce Type: new Abstract: Human disagreement is ubiquitous and well-known in labeling. However, variation in explanations, captured through token-level human rationales, remains

tutorialsarxiv-cs-cl
1 Jun 2026
Model Releases

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

DGX agent

arXiv:2511.19923v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reas

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Divergence Decoding: Inference-Time Unlearning via Auxiliary Models

DGX agent

arXiv:2605.31293v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently memorize sensitive training data thereby creating significant privacy and copyright risks. Addressing these risk

researcharxiv-cs-cl
1 Jun 2026
Tutorials

dMoE: dLLMs with Learnable Block Experts

DGX agent

arXiv:2605.30876v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance whil

tutorialsarxiv-cs-cl
1 Jun 2026
Safety

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization

DGX agent

arXiv:2605.31455v1 Announce Type: cross Abstract: Large language models are increasingly deployed in multi-turn interactive settings where users or environments can iteratively provide lightweight fee

safetyarxiv-cs-cl
1 Jun 2026
Research

Efficient Diffusion LLMs via Temporal-Spatial Parallel Decoding and Confidence Extrapolation

DGX agent

arXiv:2605.30753v1 Announce Type: new Abstract: Diffusion-based large language models (dLLMs) support parallel text generation via iterative denoising, yet inference remains latency-heavy because many

researcharxiv-cs-cl
1 Jun 2026
Model Releases

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents

DGX agent

arXiv:2605.30690v1 Announce Type: new Abstract: Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. Howeve

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents

DGX agent

arXiv:2605.30924v1 Announce Type: new Abstract: MLLM-powered embodied agents deployed in real-world environments encounter physical hazards. However, existing approaches lack explicit mechanisms for i

model-releasesarxiv-cs-cl
1 Jun 2026
Tutorials

Esoteric Language Models: A Family of Any-Order Diffusion LLMs

DGX agent

arXiv:2506.01928v4 Announce Type: replace Abstract: Diffusion-based language models offer a compelling alternative to autoregressive (AR) models by enabling parallel and controllable generation. Withi

tutorialsarxiv-cs-cl
1 Jun 2026
Model Releases

Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy

DGX agent

arXiv:2605.31506v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) is the current industry standard for grounding AI in real-world facts. Traditional retrieval methods rely on keyw

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

DGX agent

arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif

researcharxiv-cs-cl
1 Jun 2026
Research

Evidence for systematic semantic structure in individual phonemes

DGX agent

arXiv:2603.17306v3 Announce Type: replace Abstract: A foundational assumption in linguistics holds that sound-meaning relations are largely arbitrary. Here we show that this assumption fails at the le

researcharxiv-cs-cl
1 Jun 2026
Model Releases

EvoDefense: Co-Evolving Black-Box Defense with Large Language Models

DGX agent

arXiv:2605.31140v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain highly vulnerable to diverse attacks, particularly in black-box settings where the internals of target models are

model-releasesarxiv-cs-cl
1 Jun 2026
Research

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation

DGX agent

arXiv:2605.30961v1 Announce Type: new Abstract: Generating novel research ideas is fundamental to scientific progress. While Large Language Models (LLMs) show promise in assisting this process, existi

researcharxiv-cs-cl
1 Jun 2026
Model Releases

ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents

DGX agent

arXiv:2605.30712v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong capabilities in reasoning, tool use, and multi-step interaction, but they often solve tasks from scr

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

DGX agent

arXiv:2605.30947v1 Announce Type: new Abstract: LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantit

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

DGX agent

arXiv:2605.30457v1 Announce Type: cross Abstract: Regional accent classification in Brazilian Portuguese (pt-BR) suffers from the need for reliable labeling. While large self-supervised learning (SSL)

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

DGX agent

arXiv:2605.30771v1 Announce Type: new Abstract: AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence,

model-releasesarxiv-cs-cl
1 Jun 2026
Model Releases

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing

DGX agent

arXiv:2509.14221v3 Announce Type: replace-cross Abstract: Generative Engine Marketing (GEM) is an emerging ecosystem for monetizing generative engines, such as LLM-based chatbots, by seamlessly integr

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge

DGX agent

arXiv:2605.30568v1 Announce Type: new Abstract: LLM-as-a-Judge is a scalable alternative to human evaluation, yet existing rubric-based methods rely on human-annotated data such as reference answers o

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Goldfish: Monolingual Language Models for 350 Languages

DGX agent

arXiv:2408.10441v3 Announce Type: replace Abstract: For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite

model-releasesarxiv-cs-cl
1 Jun 2026
Research

GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent

DGX agent

arXiv:2603.13875v2 Announce Type: replace Abstract: Many large language model applications require conditioning on long contexts. Transformers typically support this by storing a large per-layer KV-ca

researcharxiv-cs-cl
1 Jun 2026
Research

GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs

DGX agent

arXiv:2605.31105v1 Announce Type: new Abstract: Large language models (LLMs) with extended context lengths rely on the key-value (KV) cache to support attention over prior tokens. However, maintaining

researcharxiv-cs-cl
1 Jun 2026
Research

How Much Do LLMs Know About Chinese Zero Pronouns?

DGX agent

arXiv:2605.31056v1 Announce Type: new Abstract: Zero Pronouns (ZPs) are a pervasive linguistic phenomenon in pro-drop languages such as Chinese and have long posed a challenge for natural language pro

researcharxiv-cs-cl
1 Jun 2026
Model Releases

HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds

DGX agent

arXiv:2510.15614v3 Announce Type: replace Abstract: Many scientific problems are underdetermined: multiple distinct hypotheses are equally consistent with the same observations. In such settings, effe

model-releasesarxiv-cs-cl
1 Jun 2026
Safety

IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning

DGX agent

arXiv:2602.19049v2 Announce Type: replace Abstract: Large language models increasingly rely on long chains of thought to improve accuracy, yet such gains come with substantial inference-time costs. We

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback

DGX agent

arXiv:2605.30478v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) trains language models using programmatically checkable signals such as unit-test outcomes, enab

model-releasesarxiv-cs-cl
1 Jun 2026
Research

Incremental BPE Tokenization

DGX agent

arXiv:2605.30813v1 Announce Type: new Abstract: We propose a novel algorithm for incremental Byte Pair Encoding (BPE) tokenization. The algorithm processes each input byte in worst-case O(log^2 t) tim

researcharxiv-cs-cl
1 Jun 2026
← Previous
1…6566676869…162
Next →