AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
1 Jun 2026

Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination

ResearchDGX agent

arXiv:2605.31058v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Langu

Configurable Reward Model for Balanced Safety Alignment

SafetyDGX agent

arXiv:2605.30487v1 Announce Type: new Abstract: Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

SafetyDGX agent

arXiv:2605.31073v1 Announce Type: new Abstract: Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Consolidating Rewarded Perturbations for LLM Post-Training

ResearchDGX agent

arXiv:2605.31494v1 Announce Type: new Abstract: Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by

Context-Free Recognition with Transformers

ResearchDGX agent

arXiv:2601.01754v3 Announce Type: replace-cross Abstract: Transformers excel empirically on tasks that process well-formed inputs according to some grammar, such as natural language and code. However,

Counterfactual Graph for Multi-Agent LLM Calibration

AgentsDGX agent

arXiv:2605.30653v1 Announce Type: new Abstract: Multi-agent LLM systems often treat agreement as evidence: when many agents in a panel give the same answer, that answer is assumed to be more reliable.

Cross-Lingual Steering for Figurative Language Generation

ResearchDGX agent

arXiv:2605.30443v1 Announce Type: new Abstract: Multilingual large language models can generate figurative language, but whether the internal signals driving this behavior are language-specific or reu

CSULoRA: Closest Safe Update Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2605.30640v1 Announce Type: cross Abstract: Low-rank adaptation has become a standard method for parameter-efficient fine-tuning of large language models, but even small amounts of unsafe or adv

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch

HardwareDGX agent

arXiv:2511.17826v2 Announce Type: replace-cross Abstract: Deterministic inference is increasingly critical for large language model (LLM) applications such as LLM-as-a-judge evaluation, multi-agent sy

Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection

TutorialsDGX agent

arXiv:2605.31563v1 Announce Type: new Abstract: Human disagreement is ubiquitous and well-known in labeling. However, variation in explanations, captured through token-level human rationales, remains

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

Model ReleasesDGX agent

arXiv:2511.19923v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reas

Divergence Decoding: Inference-Time Unlearning via Auxiliary Models

ResearchDGX agent

arXiv:2605.31293v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently memorize sensitive training data thereby creating significant privacy and copyright risks. Addressing these risk

dMoE: dLLMs with Learnable Block Experts

TutorialsDGX agent

arXiv:2605.30876v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive models, offering competitive performance whil

DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization

SafetyDGX agent

arXiv:2605.31455v1 Announce Type: cross Abstract: Large language models are increasingly deployed in multi-turn interactive settings where users or environments can iteratively provide lightweight fee

Efficient Diffusion LLMs via Temporal-Spatial Parallel Decoding and Confidence Extrapolation

ResearchDGX agent

arXiv:2605.30753v1 Announce Type: new Abstract: Diffusion-based large language models (dLLMs) support parallel text generation via iterative denoising, yet inference remains latency-heavy because many

ElasticMem: Latent Memory as a Learnable Resource for LLM Agents

Model ReleasesDGX agent

arXiv:2605.30690v1 Announce Type: new Abstract: Long-term memory is essential for LLM agents to reason coherently across extended interactions, personalize responses, and reuse past experience. Howeve

EMBGuard: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents

Model ReleasesDGX agent

arXiv:2605.30924v1 Announce Type: new Abstract: MLLM-powered embodied agents deployed in real-world environments encounter physical hazards. However, existing approaches lack explicit mechanisms for i

Esoteric Language Models: A Family of Any-Order Diffusion LLMs

TutorialsDGX agent

arXiv:2506.01928v4 Announce Type: replace Abstract: Diffusion-based language models offer a compelling alternative to autoregressive (AR) models by enabling parallel and controllable generation. Withi

Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy

Model ReleasesDGX agent

arXiv:2605.31506v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) is the current industry standard for grounding AI in real-world facts. Traditional retrieval methods rely on keyw

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

ResearchDGX agent

arXiv:2605.30521v1 Announce Type: new Abstract: Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classif

Evidence for systematic semantic structure in individual phonemes

ResearchDGX agent

arXiv:2603.17306v3 Announce Type: replace Abstract: A foundational assumption in linguistics holds that sound-meaning relations are largely arbitrary. Here we show that this assumption fails at the le

EvoDefense: Co-Evolving Black-Box Defense with Large Language Models

Model ReleasesDGX agent

arXiv:2605.31140v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain highly vulnerable to diverse attacks, particularly in black-box settings where the internals of target models are

EvoGens: A Population-Based Heuristic Search Framework for Scientific Idea Generation

ResearchDGX agent

arXiv:2605.30961v1 Announce Type: new Abstract: Generating novel research ideas is fundamental to scientific progress. While Large Language Models (LLMs) show promise in assisting this process, existi

ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2605.30712v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong capabilities in reasoning, tool use, and multi-step interaction, but they often solve tasks from scr

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

Model ReleasesDGX agent

arXiv:2605.30947v1 Announce Type: new Abstract: LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantit

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

ResearchDGX agent

arXiv:2605.30457v1 Announce Type: cross Abstract: Regional accent classification in Brazilian Portuguese (pt-BR) suffers from the need for reliable labeling. While large self-supervised learning (SSL)

Eywa: Provenance-Grounded Long-Term Memory for AI Agents

Model ReleasesDGX agent

arXiv:2605.30771v1 Announce Type: new Abstract: AI agents that persist across sessions need memory they can retrieve, audit, update, and erase. Existing memory systems often collapse source evidence,

GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing

Model ReleasesDGX agent

arXiv:2509.14221v3 Announce Type: replace-cross Abstract: Generative Engine Marketing (GEM) is an emerging ecosystem for monetizing generative engines, such as LLM-based chatbots, by seamlessly integr

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge

ResearchDGX agent

arXiv:2605.30568v1 Announce Type: new Abstract: LLM-as-a-Judge is a scalable alternative to human evaluation, yet existing rubric-based methods rely on human-annotated data such as reference answers o

Goldfish: Monolingual Language Models for 350 Languages

Model ReleasesDGX agent

arXiv:2408.10441v3 Announce Type: replace Abstract: For many low-resource languages, the only available language models are large multilingual models trained on many languages simultaneously. Despite

GradMem: Learning to Write Context into Memory with Test-Time Gradient Descent

ResearchDGX agent

arXiv:2603.13875v2 Announce Type: replace Abstract: Many large language model applications require conditioning on long contexts. Transformers typically support this by storing a large per-layer KV-ca

GRKV: Global Regression for Training-Free KV Cache Compression in Long-Context LLMs

ResearchDGX agent

arXiv:2605.31105v1 Announce Type: new Abstract: Large language models (LLMs) with extended context lengths rely on the key-value (KV) cache to support attention over prior tokens. However, maintaining

How Much Do LLMs Know About Chinese Zero Pronouns?

ResearchDGX agent

arXiv:2605.31056v1 Announce Type: new Abstract: Zero Pronouns (ZPs) are a pervasive linguistic phenomenon in pro-drop languages such as Chinese and have long posed a challenge for natural language pro

HypoSpace: A Diagnostic Benchmark for Set-Valued Hypothesis Generation under Underdetermination and Sublinear Coverage Bounds

Model ReleasesDGX agent

arXiv:2510.15614v3 Announce Type: replace Abstract: Many scientific problems are underdetermined: multiple distinct hypotheses are equally consistent with the same observations. In such settings, effe

IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning

SafetyDGX agent

arXiv:2602.19049v2 Announce Type: replace Abstract: Large language models increasingly rely on long chains of thought to improve accuracy, yet such gains come with substantial inference-time costs. We

Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback

Model ReleasesDGX agent

arXiv:2605.30478v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) trains language models using programmatically checkable signals such as unit-test outcomes, enab

Incremental BPE Tokenization

ResearchDGX agent

arXiv:2605.30813v1 Announce Type: new Abstract: We propose a novel algorithm for incremental Byte Pair Encoding (BPE) tokenization. The algorithm processes each input byte in worst-case O(log^2 t) tim

'Intelegi Romaneste?'' A Recipe for Romanian Vision-Language Models

ResearchDGX agent

arXiv:2605.31401v1 Announce Type: new Abstract: Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languag

Knowledge Boundary Probing and Demand-Guided Intervention for LLM-Based Power System Code Generation

Model ReleasesDGX agent

arXiv:2605.31478v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate power-system analysis, but many utilities and energy-research labs require on-premise s

Knowledge Graph-Enhanced Zero-Shot Topic Classification: A Multi-Strategy Comparative Study

ResearchDGX agent

arXiv:2605.30465v1 Announce Type: new Abstract: Multi-label topic classification without labeled training data is a challenging task, specially when documents contain complex relational information. W

Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task

ResearchDGX agent

arXiv:2605.31480v1 Announce Type: new Abstract: Do neural models, such as Large Language Models, genuinely acquire compositional abilities for interpretation of natural language? When we talk about se

Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization

ResearchDGX agent

arXiv:2605.31312v1 Announce Type: cross Abstract: Multimodal hallucination remains a persistent challenge for Vision-Language Models (VLMs). Standard textual Direct Preference Optimization (DPO) often

Learning Whom to Trust: Market-Feedback Adaptive Retrieval for Frozen LLMs in Event-Driven Financial RAG

Model ReleasesDGX agent

arXiv:2605.31201v1 Announce Type: new Abstract: Financial retrieval-augmented generation (RAG) systems typically rank evidence by textual relevance, but in financial markets the useful evidence source

Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

ResearchDGX agent

arXiv:2605.30501v1 Announce Type: new Abstract: Watermarking embeds statistical signatures in AI-generated text for detection and attribution. We reveal a fundamental vulnerability: when users access

LLM Anonymization Against Agentic Re-Identificatio

AgentsDGX agent

arXiv:2605.30848v1 Announce Type: cross Abstract: Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-ident

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

SafetyDGX agent

arXiv:2605.31381v1 Announce Type: new Abstract: We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that

LocalSUG: City-Preference-Enhanced LLM for Query Suggestion in Local-Life Services

Local AiDGX agent

arXiv:2603.04946v2 Announce Type: replace Abstract: In local-life service platforms, query suggestion reduces user effort by generating candidate queries from input prefixes. Traditional multi-stage s

MAAT: Multi-phase Adapter-Aware Targeted Unlearning

Model ReleasesDGX agent

arXiv:2605.30514v1 Announce Type: cross Abstract: Machine unlearning evaluation is structurally skewed: Why-type questions, which probe causal and relational knowledge, comprise less than 0.06% of Cou

MADS: Model-Aware Diverse Core Set Selection for Instruction Tuning

Model ReleasesDGX agent

arXiv:2605.30857v1 Announce Type: new Abstract: Instruction fine-tuning is employed to enhance the instruction-following ability of large language models (LLMs). As the amount of instruction fine-tuni

Measuring, Localizing, and Ablating Alignment Signatures in LLMs

Local AiDGX agent

arXiv:2605.30526v1 Announce Type: cross Abstract: Aligned language models often exhibit a recognizable AI-like style, yet its connection to post-training and internal representations remains poorly un

Mellum2 Technical Report

Model ReleasesDGX agent

arXiv:2605.31268v1 Announce Type: new Abstract: We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-p

Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning

Local AiDGX agent

arXiv:2602.13069v2 Announce Type: replace-cross Abstract: On-device fine-tuning enables privacy-preserving personalization of large language models, but mobile devices impose severe memory constraints

MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

Model ReleasesDGX agent

arXiv:2605.30931v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to susta

MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2605.31010v1 Announce Type: new Abstract: Retrieval-augmented generation is intensively studied to ground large language models on external evidence. However, retrieving from a unified knowledge

MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents

Model ReleasesDGX agent

arXiv:2605.30727v1 Announce Type: new Abstract: Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent's external q

Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely

AgentsDGX agent

arXiv:2605.31387v1 Announce Type: new Abstract: Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected

Multilingual and Cross-Lingual Citation Needed Detection on Wikipedia for Lower-Resource Languages

ResearchDGX agent

arXiv:2605.31136v1 Announce Type: new Abstract: In automated fact-checking (AFC), check-worthiness detection identifies claims requiring verification based on domain-specific criteria. On Wikipedia, t

NeUQI: Near-Optimal Uniform Quantization Parameter Initialization for Low-Bit LLMs

Model ReleasesDGX agent

arXiv:2505.17595v4 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve impressive performance across domains but face significant challenges when deployed on consumer-grade GPU

Neuron-Level Interventions for Gendered and Gender-Neutral Generation in Language Models

SafetyDGX agent

arXiv:2605.30717v1 Announce Type: new Abstract: Language models (LMs) can produce gendered language and stereotypes even when given neutral prompts. Most prior work on gender bias in LMs primarily exa

On the 'Induction Bias' in Sequence Models

SafetyDGX agent

arXiv:2602.18333v2 Announce Type: replace-cross Abstract: Despite the remarkable practical success of transformer-based language models, recent work has raised concerns about their ability to perform

← Previous
1…5152535455…129
Next →