AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Safety

HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems

DGX agent

arXiv:2606.01779v1 Announce Type: new Abstract: LLM agents are increasingly expected to operate across heterogeneous task regimes that require distinct execution paradigms. This challenges fixed agent

safetyarxiv-cs-cl
2 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

HERO'S JOURNEY: Testing Complex Rule Induction with Text Games

DGX agent

arXiv:2606.02556v1 Announce Type: new Abstract: We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations an

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression

DGX agent

arXiv:2606.01934v1 Announce Type: cross Abstract: Large language models achieve remarkable performance via extended chain-of-thought (CoT) reasoning, yet this lengthy process incurs substantial infere

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

How AI Fails: An Interactive Pedagogical Tool for Demonstrating Dialectal Bias in Automated Toxicity Models

DGX agent

arXiv:2511.06676v3 Announce Type: replace Abstract: Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that 'the AI is biased'. While this is often said jokingly

model-releasesarxiv-cs-cl
2 Jun 2026
Research

How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings

DGX agent

arXiv:2606.00356v1 Announce Type: new Abstract: Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary

researcharxiv-cs-cl
2 Jun 2026
Model Releases

How to Correctly Report LLM-as-a-Judge Evaluations

DGX agent

arXiv:2511.21140v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensiti

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

HypothesisMed: Inference-Time Answer Fusion and Structured Hypothesis-Space Reporting for Biomedical Question Answering

DGX agent

arXiv:2606.00971v1 Announce Type: new Abstract: Biomedical question answering with large language models is commonly evaluated using answer accuracy, but answer accuracy alone does not indicate whethe

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

'I Strongly Suspect This Website Is a Scam': Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents

DGX agent

arXiv:2606.00497v1 Announce Type: cross Abstract: Deceptive web content, widely instantiated across the internet and commonly known as extit{social-engineering attacks}, manipulates autonomous web age

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications

DGX agent

arXiv:2606.00750v1 Announce Type: new Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing

model-releasesarxiv-cs-cl
2 Jun 2026
Research

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

DGX agent

arXiv:2606.00875v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks involving creative problem solving and idea generation. However, there is a lack of consens

researcharxiv-cs-cl
2 Jun 2026
Applications

InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models

DGX agent

arXiv:2606.02161v1 Announce Type: cross Abstract: Video Large Language Models (Video-LLMs) achieve strong performance in video understanding, but their excessive visual tokens bring substantial comput

applicationsarxiv-cs-cl
2 Jun 2026
Safety

Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning

DGX agent

arXiv:2606.00755v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in whic

safetyarxiv-cs-cl
2 Jun 2026
Research

Interpreto: An Explainability Library for Transformers

DGX agent

arXiv:2512.09730v3 Announce Type: replace Abstract: Interpreto is an open-source Python library for interpreting HuggingFace language models, from early BERT variants to LLMs. It provides two compleme

researcharxiv-cs-cl
2 Jun 2026
Model Releases

Investigating and Alleviating Harm Amplification in LLM Interactions

DGX agent

arXiv:2606.02423v1 Announce Type: new Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve ha

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

DGX agent

arXiv:2606.02404v1 Announce Type: new Abstract: Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, b

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices

DGX agent

arXiv:2601.21579v2 Announce Type: replace Abstract: The success of Hyper-Connections (HC) in neural networks (NN) has also highlighted issues related to training instability and restricted scalability

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

LaSR: Context-Aware Speech Recognition via Latent Reasoning

DGX agent

arXiv:2606.00507v1 Announce Type: new Abstract: Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their co

model-releasesarxiv-cs-cl
2 Jun 2026
Tutorials

Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning

DGX agent

arXiv:2511.07910v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve excellent performance in natural language reasoning tasks through pre-training on vast unstructured text, enabl

tutorialsarxiv-cs-cl
2 Jun 2026
Research

Learning from Saturated Data: Signals Beyond Correctness for LLM Training

DGX agent

arXiv:2606.01436v1 Announce Type: new Abstract: The growing capabilities of large language models (LLMs) have led to the saturation of many benchmarks and training datasets used to improve them. Motiv

researcharxiv-cs-cl
2 Jun 2026
Agents

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

DGX agent

arXiv:2602.03619v2 Announce Type: replace Abstract: Nowadays, developing reliable DeepResearch-style long-form report generation remains challenging, as training and evaluation lack verifiable reward

agentsarxiv-cs-cl
2 Jun 2026
Model Releases

Learning to Retrieve: Dual-Level Long-Term Memory for Text-to-SQL Agents

DGX agent

arXiv:2606.00547v1 Announce Type: new Abstract: Interactive text-to-SQL agents solve database tasks through multi-turn interactions involving schema exploration, query execution, feedback interpretati

model-releasesarxiv-cs-cl
2 Jun 2026
Research

Lessons from the Trenches on Reproducible Evaluation of Language Models

DGX agent

arXiv:2405.14782v3 Announce Type: replace Abstract: Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivi

researcharxiv-cs-cl
2 Jun 2026
Research

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

DGX agent

arXiv:2602.23881v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate t

researcharxiv-cs-cl
2 Jun 2026
Safety

LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation

DGX agent

arXiv:2603.09403v2 Announce Type: replace Abstract: Validating evaluation metrics for NLG typically relies on expensive and time-consuming human annotations, which predominantly exist only for English

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection

DGX agent

arXiv:2606.00684v1 Announce Type: cross Abstract: We address the problem of out-of-distribution (OOD) detection for target observations embedded in a subspace of the high dimensional data space. Using

model-releasesarxiv-cs-cl
2 Jun 2026
Applications

LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning

DGX agent

arXiv:2606.01336v1 Announce Type: new Abstract: As real-world applications increasingly require processing inputs of 100k+ tokens, the gap between context length and inference efficiency has become a

applicationsarxiv-cs-cl
2 Jun 2026
Safety

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

DGX agent

arXiv:2606.00975v1 Announce Type: new Abstract: LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delu

safetyarxiv-cs-cl
2 Jun 2026
Research

M^3 Scaling Law: Optimizing Multi-Epoch, Multi-Lingual, and Multi-Stage Training for Low-Resource Language Models

DGX agent

arXiv:2410.12325v2 Announce Type: replace Abstract: In this paper, we study a fundamental design problem in pretraining Large Language Models (LLMs) for low-resource language regimes. Existing works a

researcharxiv-cs-cl
2 Jun 2026
Research

Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling

DGX agent

arXiv:2606.02004v1 Announce Type: new Abstract: Consumer-price measurement increasingly draws on alternative data sources -- scanner, web-scraped, and transaction/receipt data. A recurring obstacle is

researcharxiv-cs-cl
2 Jun 2026
Research

Malaysian English News Decoded: A Linguistic Resource for Named Entity and Relation Extraction

DGX agent

arXiv:2402.14521v2 Announce Type: replace Abstract: Standard English and Malaysian English exhibit notable differences, posing challenges for natural language processing (NLP) tasks on Malaysian Engli

researcharxiv-cs-cl
2 Jun 2026
Model Releases

MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation

DGX agent

arXiv:2505.18614v5 Announce Type: replace Abstract: Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style. In animated mu

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

DGX agent

arXiv:2606.01914v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly atten

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning

DGX agent

arXiv:2606.01301v1 Announce Type: new Abstract: Hallucinations in medical large language models (LLMs) pose serious risks for clinical decision support, particularly when models must reason over compl

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment

DGX agent

arXiv:2603.20884v2 Announce Type: replace Abstract: To alleviate the heavy burden of paper screening, researchers increasingly rely on existing AI agents, such as AI reviewers or DeepResearch, for pap

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

DGX agent

arXiv:2602.03318v3 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling-a slow and fragile process ill-suited to novel scenarios. While large language models (LLM

agentsarxiv-cs-cl
2 Jun 2026
Safety

Mitigating Bias in Locally Constrained Decoding via Tractable Proposals

DGX agent

arXiv:2606.01926v1 Announce Type: new Abstract: Generations from large language models often fail to conform to desired constraints such as JSON schema. Existing locally constrained decoding (LCD) app

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Model-Based Quality Assessment for Massively Multilingual Parallel Data

DGX agent

arXiv:2606.00285v1 Announce Type: new Abstract: Large-scale multilingual bitext often contains two distinct problems: non-parallel sentence pairs and low-quality translations. We decompose model-based

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

Modeling Distinct Human Interaction in Web Agents

DGX agent

arXiv:2602.17588v3 Announce Type: replace Abstract: Despite rapid progress in autonomous web agents, human involvement remains essential for shaping preferences and correcting agent behavior as tasks

agentsarxiv-cs-cl
2 Jun 2026
Model Releases

Momento: Evaluating Persistent Memory and Reasoning with Multi-Session Agentic Conversations

DGX agent

arXiv:2606.00832v1 Announce Type: new Abstract: Recent advances in agentic AI have enabled agents to complete complex tasks through tool use, reasoning, and multi-step planning. Yet existing benchmark

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Multi-Agent Computer Use

DGX agent

arXiv:2606.01533v1 Announce Type: cross Abstract: Computer use agents (CUAs) today are primarily deployed as single serial agents. This setup is suboptimal for complex long-horizon tasks that benefit

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Navigating the Reality Gap: On-Device Continual Adaptation of ASR for Clinical Telephony

DGX agent

arXiv:2512.16401v5 Announce Type: replace Abstract: Automatic Speech Recognition (ASR) can significantly reduce documentation burden in clinical workflows, but standard models degrade sharply in real-

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

NormEval: A Unified Multi-Metric Framework for Evaluating Semantic Fidelity in Text Normalization

DGX agent

arXiv:2511.20409v2 Announce Type: replace Abstract: Text normalization methods such as stemming and lemmatization are fundamental components of NLP pipelines. As new normalization tools are developed

safetyarxiv-cs-cl
2 Jun 2026
Research

Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales

DGX agent

arXiv:2606.01148v1 Announce Type: new Abstract: Natural-language explanations are often treated as a unified interface for understanding model behavior, but different explanation sources may support s

researcharxiv-cs-cl
2 Jun 2026
Agents

Not All Flips Are Conformity: Decomposing Stance Convergence in Multi-Agent LLM Debate

DGX agent

arXiv:2606.00820v1 Announce Type: new Abstract: Multi-agent debate (MAD) is a promising strategy for improving LLM reasoning, but when agents converge on a shared answer, it is unclear whether that co

agentsarxiv-cs-cl
2 Jun 2026
Research

Not What, But How: A Communicative Audit of LLM Response Framing

DGX agent

arXiv:2606.02493v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses

researcharxiv-cs-cl
2 Jun 2026
Model Releases

OARelatedWork: A Large-Scale Dataset of Related Work Sections with Full-texts from Open Access Sources

DGX agent

arXiv:2405.01930v2 Announce Type: replace Abstract: This paper introduces OARelatedWork: a dataset for related work generation from open-access sources. It is the first large-scale multi-document summ

model-releasesarxiv-cs-cl
2 Jun 2026
Research

OCC-RAG: Optimal Cognitive Core for Faithful Question Answering

DGX agent

arXiv:2606.00683v1 Announce Type: new Abstract: Recent progress in the development of language models has been defined by scale, with each generation absorbing more of the world's knowledge into its w

researcharxiv-cs-cl
2 Jun 2026
Model Releases

OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification

DGX agent

arXiv:2606.01476v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitig

model-releasesarxiv-cs-cl
2 Jun 2026
← Previous
1…6162636465…161
Next →