AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Model Releases

Functional Emotions or Situational Contexts? A Discriminating Test from the Mythos Preview System Card

DGX agent

arXiv:2604.13466v1 Announce Type: cross Abstract: The Claude Mythos Preview system card deploys emotion vectors, sparse autoencoder (SAE) features, and activation verbalisers to study model internals

model-releasesarxiv-cs-cl
16 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus

DGX agent

arXiv:2604.13288v1 Announce Type: new Abstract: We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruvian Constitution using three state-of-the-art text-t

applicationsarxiv-cs-cl
16 Apr 2026
Model Releases

Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs

DGX agent

arXiv:2604.13258v1 Announce Type: new Abstract: Attribution methods seek to explain language model predictions by quantifying the contribution of input tokens to generated outputs. However, most exist

model-releasesarxiv-cs-cl
16 Apr 2026
Research

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

DGX agent

arXiv:2604.13977v1 Announce Type: new Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing stra

researcharxiv-cs-cl
16 Apr 2026
Research

(How) Learning Rates Regulate Catastrophic Overtraining

DGX agent

arXiv:2604.13627v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) is a common first stage of LLM post-training, teaching the model to follow instructions and shaping its behavior as a hel

researcharxiv-cs-cl
16 Apr 2026
Model Releases

Hybrid Retrieval for COVID-19 Literature: Comparing Rank Fusion and Projection Fusion with Diversity Reranking

DGX agent

arXiv:2604.13728v1 Announce Type: cross Abstract: We present a hybrid retrieval system for COVID-19 scientific literature, evaluated on the TREC-COVID benchmark (171,332 papers, 50 expert queries). Th

model-releasesarxiv-cs-cl
16 Apr 2026
Research

Indexing Multimodal Language Models for Large-scale Image Retrieval

DGX agent

arXiv:2604.13268v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong cross-modal reasoning capabilities, yet their potential for vision-only tasks remain

researcharxiv-cs-cl
16 Apr 2026
Model Releases

IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages

DGX agent

arXiv:2604.13686v1 Announce Type: new Abstract: While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

DGX agent

arXiv:2604.13201v1 Announce Type: new Abstract: Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks

model-releasesarxiv-cs-cl
16 Apr 2026
Tutorials

Interpretable Stylistic Variation in Human and LLM Writing Across Genres, Models, and Decoding Strategies

DGX agent

arXiv:2604.14111v1 Announce Type: new Abstract: Large Language Models (LLMs) are now capable of generating highly fluent, human-like text. They enable many applications, but also raise concerns such a

tutorialsarxiv-cs-cl
16 Apr 2026
Applications

IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages

DGX agent

arXiv:2604.13078v1 Announce Type: new Abstract: The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts

applicationsarxiv-cs-cl
16 Apr 2026
Research

Just Use XML: Revisiting Joint Translation and Label Projection

DGX agent

arXiv:2603.12021v2 Announce Type: replace Abstract: Label projection is an effective technique for cross-lingual transfer, extending span-annotated datasets from a high-resource language to low-resour

researcharxiv-cs-cl
16 Apr 2026
Model Releases

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

DGX agent

arXiv:2604.13058v1 Announce Type: new Abstract: We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,46

model-releasesarxiv-cs-cl
16 Apr 2026
Tutorials

Kwame 2.0: Human-in-the-Loop Generative AI Teaching Assistant for Large Scale Online Coding Education in Africa

DGX agent

arXiv:2603.29159v2 Announce Type: replace Abstract: Providing timely and accurate learning support in large-scale online coding courses is challenging, particularly in resource-constrained contexts. W

tutorialsarxiv-cs-cl
16 Apr 2026
Model Releases

L2D-Clinical: Learning to Defer for Adaptive Model Selection in Clinical Text Classification

DGX agent

arXiv:2604.13285v1 Announce Type: new Abstract: Clinical text classification requires choosing between specialized fine-tuned models (BERT variants) and general-purpose large language models (LLMs), y

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Language steering in latent space to mitigate unintended code-switching

DGX agent

arXiv:2510.13849v3 Announce Type: replace Abstract: Multilingual Large Language Models (LLMs) often exhibit hallucinations such as unintended code-switching, reducing reliability in downstream tasks.

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models

DGX agent

arXiv:2511.11334v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Learning the Cue or Learning the Word? Analyzing Generalization in Metaphor Detection for Verbs

DGX agent

arXiv:2604.13713v1 Announce Type: new Abstract: Metaphor detection models achieve strong benchmark performance, yet it remains unclear whether this reflects transferable generalization or lexical memo

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Leveraging LLM-GNN Integration for Open-World Question Answering over Knowledge Graphs

DGX agent

arXiv:2604.13979v1 Announce Type: new Abstract: Open-world Question Answering (OW-QA) over knowledge graphs (KGs) aims to answer questions over incomplete or evolving KGs. Traditional KGQA assumes a c

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

DGX agent

arXiv:2604.13072v1 Announce Type: new Abstract: LLM-based agents are increasingly expected to handle real-world assistant tasks, yet existing benchmarks typically evaluate them under isolated sources

model-releasesarxiv-cs-cl
16 Apr 2026
Applications

Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning

DGX agent

arXiv:2601.02902v2 Announce Type: replace-cross Abstract: Symbolic logical reasoning is a critical yet underexplored capability of large language models (LLMs), providing reliable and verifiable decis

applicationsarxiv-cs-cl
16 Apr 2026
Model Releases

Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data

DGX agent

arXiv:2604.13066v1 Announce Type: new Abstract: In-context learning has established itself as an important learning paradigm for Large Language Models (LLMs). In this paper, we demonstrate that LLMs c

model-releasesarxiv-cs-cl
16 Apr 2026
Applications

Mathematical Reasoning Enhanced LLM for Formula Derivation: A Case Study on Fiber NLI Modellin

DGX agent

arXiv:2604.13062v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have demonstrated strong capabilities in code generation and text synthesis, yet their potential for sym

applicationsarxiv-cs-cl
16 Apr 2026
Model Releases

MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging

DGX agent

arXiv:2604.13756v1 Announce Type: new Abstract: The potential of Multimodal Large Language Models (MLLMs) in domain of medical imaging raise the demands of systematic and rigorous evaluation framework

model-releasesarxiv-cs-cl
16 Apr 2026
Applications

Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents

DGX agent

arXiv:2604.14004v1 Announce Type: cross Abstract: Memory-based self-evolution has emerged as a promising paradigm for coding agents. However, existing approaches typically restrict memory utilization

applicationsarxiv-cs-cl
16 Apr 2026
Agents

Memp: Exploring Agent Procedural Memory

DGX agent

arXiv:2508.06433v4 Announce Type: replace Abstract: Large Language Models (LLMs) based agents excel at diverse tasks, yet they suffer from brittle procedural memory that is manually engineered or enta

agentsarxiv-cs-cl
16 Apr 2026
Model Releases

MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments

DGX agent

arXiv:2604.13418v1 Announce Type: new Abstract: Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web re

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates

DGX agent

arXiv:2512.04844v2 Announce Type: replace Abstract: Expanding the linguistic diversity of instruct large language models (LLMs) is crucial for global accessibility but is often hindered by the relianc

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

MM-Doc-R1: Training Agents for Long Document Visual Question Answering through Multi-turn Reinforcement Learning

DGX agent

arXiv:2604.13579v1 Announce Type: new Abstract: Conventional Retrieval-Augmented Generation (RAG) systems often struggle with complex multi-hop queries over long documents due to their single-pass ret

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

DGX agent

arXiv:2505.07591v2 Announce Type: replace Abstract: Instruction following refers to the ability of large language models (LLMs) to generate outputs that satisfy all specified constraints. Existing res

model-releasesarxiv-cs-cl
16 Apr 2026
Local Ai

MUSE: Multi-Domain Chinese User Simulation via Self-Evolving Profiles and Rubric-Guided Alignment

DGX agent

arXiv:2604.13828v1 Announce Type: new Abstract: User simulators are essential for the scalable training and evaluation of interactive AI systems. However, existing approaches often rely on shallow use

local-aiarxiv-cs-cl
16 Apr 2026
Research

Native Hybrid Attention for Efficient Sequence Modeling

DGX agent

arXiv:2510.07019v3 Announce Type: replace Abstract: Transformers excel at sequence modeling but face quadratic complexity, while linear attention offers improved efficiency but often compromises recal

researcharxiv-cs-cl
16 Apr 2026
Research

Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Models

DGX agent

arXiv:2601.11340v2 Announce Type: replace Abstract: Chain-of-Thought reasoning has significantly enhanced the problem-solving capabilities of Large Language Models. Unfortunately, current models gener

researcharxiv-cs-cl
16 Apr 2026
Safety

Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning

DGX agent

arXiv:2506.08125v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong reasoning abilities but often produce unnecessarily long explanations that reduce efficiency. Althoug

safetyarxiv-cs-cl
16 Apr 2026
Research

OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs

DGX agent

arXiv:2604.13073v1 Announce Type: new Abstract: Modern multimodal large language models (MLLMs) generate fluent responses from interleaved text, image, audio, and video inputs. However, identifying wh

researcharxiv-cs-cl
16 Apr 2026
Model Releases

Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning

DGX agent

arXiv:2604.14010v1 Announce Type: cross Abstract: Supervised Fine-Tuning (SFT) of large language models often suffers from task interference and catastrophic forgetting. Recent approaches alleviate th

model-releasesarxiv-cs-cl
16 Apr 2026
Research

ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian

DGX agent

arXiv:2511.01619v2 Announce Type: replace Abstract: ParlaSpeech is a collection of spoken parliamentary corpora currently spanning four Slavic languages - Croatian, Czech, Polish and Serbian - all tog

researcharxiv-cs-cl
16 Apr 2026
Model Releases

Peer-Predictive Self-Training for Language Model Reasoning

DGX agent

arXiv:2604.13356v1 Announce Type: new Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Trai

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

PersonaVLM: Long-Term Personalized Multimodal LLMs

DGX agent

arXiv:2604.13074v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual pr

model-releasesarxiv-cs-cl
16 Apr 2026
Agents

pi-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data

DGX agent

arXiv:2604.14054v1 Announce Type: cross Abstract: Deep search agents have emerged as a promising paradigm for addressing complex information-seeking tasks, but their training remains challenging due t

agentsarxiv-cs-cl
16 Apr 2026
Research

QuantileMark: A Message-Symmetric Multi-bit Watermark for LLMs

DGX agent

arXiv:2604.13786v1 Announce Type: new Abstract: As large language models become standard backends for content generation, practical provenance increasingly requires multi-bit watermarking. In provider

researcharxiv-cs-cl
16 Apr 2026
Agents

RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows

DGX agent

arXiv:2509.20490v4 Announce Type: replace-cross Abstract: Agentic systems offer a potential path to solve complex clinical tasks through collaboration among specialized agents, augmented by tool use a

agentsarxiv-cs-cl
16 Apr 2026
Model Releases

RAG or Learning? Understanding the Limits of LLM Adaptation under Continuous Knowledge Drift in the Real World

DGX agent

arXiv:2604.05096v2 Announce Type: replace Abstract: Large language models (LLMs) acquire most of their knowledge during pretraining, which ties them to a fixed snapshot of the world and makes adaptati

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Red Skills or Blue Skills? A Dive Into Skills Published on ClawHub

DGX agent

arXiv:2604.13064v1 Announce Type: new Abstract: Skill ecosystems have emerged as an increasingly important layer in Large Language Model (LLM) agent systems, enabling reusable task packaging, public d

model-releasesarxiv-cs-cl
16 Apr 2026
Safety

Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

DGX agent

arXiv:2601.03027v3 Announce Type: replace Abstract: Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce halluc

safetyarxiv-cs-cl
16 Apr 2026
Agents

Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution

DGX agent

arXiv:2512.10696v2 Announce Type: replace-cross Abstract: Procedural memory enables large language model (LLM) agents to internalize 'how-to' knowledge, theoretically reducing redundant trial-and-erro

agentsarxiv-cs-cl
16 Apr 2026
Model Releases

Reward Design for Physical Reasoning in Vision-Language Models

DGX agent

arXiv:2604.13993v1 Announce Type: cross Abstract: Physical reasoning over visual inputs demands tight integration of visual perception, domain knowledge, and multi-step symbolic inference. Yet even st

model-releasesarxiv-cs-cl
16 Apr 2026
Local Ai

Rhetorical Questions in LLM Representations: A Linear Probing Study

DGX agent

arXiv:2604.14128v1 Announce Type: new Abstract: Rhetorical questions are asked not to seek information but to persuade or signal stance. How large language models internally represent them remains unc

local-aiarxiv-cs-cl
16 Apr 2026
← Previous
1…146147148149150…160
Next →