AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Research

From Global to Local: Learning Context-Aware Graph Representations for Document Classification and Summarization

DGX agent

arXiv:2603.00021v2 Announce Type: replace Abstract: Recent NLP systems commonly represent documents as linear token sequences. Although this captures sequential order, it can hinder modeling long-rang

researcharxiv-cs-cl
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication

DGX agent

arXiv:2606.01136v1 Announce Type: new Abstract: Single-score translation metrics can conflate legitimate variation with error, a problem especially acute for classical languages where multiple defensi

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

From Unfamiliar to Familiar: Detecting Pre-training Data via Gradient Deviations in Large Language Models

DGX agent

arXiv:2603.04828v2 Announce Type: replace Abstract: Pre-training data detection for LLMs is essential for addressing copyright concerns and mitigating benchmark contamination. Existing methods mainly

model-releasesarxiv-cs-cl
2 Jun 2026
Research

GateKD: Confidence-Gated Closed-Loop Distillation for Robust Reasoning

DGX agent

arXiv:2605.13136v2 Announce Type: replace Abstract: Distilling multi-step reasoning abilities from large language models (LLMs) into compact student models remains challenging due to noisy rationales,

researcharxiv-cs-cl
2 Jun 2026
Research

GeistBERT: Breathing Life into German NLP

DGX agent

arXiv:2506.11903v5 Announce Type: replace Abstract: Advances in transformer-based language models have highlighted the benefits of language-specific pre-training on high-quality corpora. In this conte

researcharxiv-cs-cl
2 Jun 2026
Research

Geometric Latent Reasoning Induces Shorter Generations in LLMs

DGX agent

arXiv:2606.02248v1 Announce Type: new Abstract: Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive,

researcharxiv-cs-cl
2 Jun 2026
Model Releases

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

DGX agent

arXiv:2510.24081v2 Announce Type: replace Abstract: To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and

model-releasesarxiv-cs-cl
2 Jun 2026
Research

GottBERT: a pure German Language Model

DGX agent

arXiv:2012.02110v2 Announce Type: replace Abstract: Pre-trained language models have significantly advanced natural language processing (NLP), especially with the introduction of BERT and its optimize

researcharxiv-cs-cl
2 Jun 2026
Applications

Graph-Augmented Retrieval for Cross-Entity Financial Sentiment Analysis: A Comparative Study

DGX agent

arXiv:2606.00062v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become foundational for grounding large language models in domain-specific corpora, yet conventional vector-bas

applicationsarxiv-cs-cl
2 Jun 2026
Safety

Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation

DGX agent

arXiv:2510.18439v3 Announce Type: replace Abstract: Hallucination, where models generate fluent text unsupported by visual evidence, remains a major flaw in vision-language models and is particularly

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

HalleluBERT: Let Every Token That Has Meaning Bear Its Weight

DGX agent

arXiv:2510.21372v2 Announce Type: replace Abstract: Transformer-based models have advanced NLP, yet Hebrew still lacks a RoBERTa encoder that is trained at scale and released in both base and large va

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems

DGX agent

arXiv:2606.01779v1 Announce Type: new Abstract: LLM agents are increasingly expected to operate across heterogeneous task regimes that require distinct execution paradigms. This challenges fixed agent

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

HERO'S JOURNEY: Testing Complex Rule Induction with Text Games

DGX agent

arXiv:2606.02556v1 Announce Type: new Abstract: We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations an

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression

DGX agent

arXiv:2606.01934v1 Announce Type: cross Abstract: Large language models achieve remarkable performance via extended chain-of-thought (CoT) reasoning, yet this lengthy process incurs substantial infere

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

How AI Fails: An Interactive Pedagogical Tool for Demonstrating Dialectal Bias in Automated Toxicity Models

DGX agent

arXiv:2511.06676v3 Announce Type: replace Abstract: Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that 'the AI is biased'. While this is often said jokingly

model-releasesarxiv-cs-cl
2 Jun 2026
Research

How Far Do Auto-Interpretation Labels Generalize: A Controlled Study Across Languages, Scripts, and Rewordings

DGX agent

arXiv:2606.00356v1 Announce Type: new Abstract: Sparse autoencoder (SAE) features are increasingly used to interpret language models, with auto-generated natural-language labels serving as the primary

researcharxiv-cs-cl
2 Jun 2026
Model Releases

How to Correctly Report LLM-as-a-Judge Evaluations

DGX agent

arXiv:2511.21140v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators. However, imperfect sensiti

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

HypothesisMed: Inference-Time Answer Fusion and Structured Hypothesis-Space Reporting for Biomedical Question Answering

DGX agent

arXiv:2606.00971v1 Announce Type: new Abstract: Biomedical question answering with large language models is commonly evaluated using answer accuracy, but answer accuracy alone does not indicate whethe

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

'I Strongly Suspect This Website Is a Scam': Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents

DGX agent

arXiv:2606.00497v1 Announce Type: cross Abstract: Deceptive web content, widely instantiated across the internet and commonly known as extit{social-engineering attacks}, manipulates autonomous web age

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications

DGX agent

arXiv:2606.00750v1 Announce Type: new Abstract: Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing

model-releasesarxiv-cs-cl
2 Jun 2026
Research

IDEAFix: Evaluation Framework for Creative Defixation Prompting in LLMs

DGX agent

arXiv:2606.00875v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for tasks involving creative problem solving and idea generation. However, there is a lack of consens

researcharxiv-cs-cl
2 Jun 2026
Applications

InfoMerge: Information-aware Token Compression for Efficient Video Large Language Models

DGX agent

arXiv:2606.02161v1 Announce Type: cross Abstract: Video Large Language Models (Video-LLMs) achieve strong performance in video understanding, but their excessive visual tokens bring substantial comput

applicationsarxiv-cs-cl
2 Jun 2026
Safety

Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning

DGX agent

arXiv:2606.00755v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in whic

safetyarxiv-cs-cl
2 Jun 2026
Research

Interpreto: An Explainability Library for Transformers

DGX agent

arXiv:2512.09730v3 Announce Type: replace Abstract: Interpreto is an open-source Python library for interpreting HuggingFace language models, from early BERT variants to LLMs. It provides two compleme

researcharxiv-cs-cl
2 Jun 2026
Model Releases

Investigating and Alleviating Harm Amplification in LLM Interactions

DGX agent

arXiv:2606.02423v1 Announce Type: new Abstract: Large language models (LLMs) can serve as helpful assistants, yet they can equally function as harm amplifiers that enable malicious users to achieve ha

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

DGX agent

arXiv:2606.02404v1 Announce Type: new Abstract: Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, b

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices

DGX agent

arXiv:2601.21579v2 Announce Type: replace Abstract: The success of Hyper-Connections (HC) in neural networks (NN) has also highlighted issues related to training instability and restricted scalability

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

LaSR: Context-Aware Speech Recognition via Latent Reasoning

DGX agent

arXiv:2606.00507v1 Announce Type: new Abstract: Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their co

model-releasesarxiv-cs-cl
2 Jun 2026
Tutorials

Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning

DGX agent

arXiv:2511.07910v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve excellent performance in natural language reasoning tasks through pre-training on vast unstructured text, enabl

tutorialsarxiv-cs-cl
2 Jun 2026
Research

Learning from Saturated Data: Signals Beyond Correctness for LLM Training

DGX agent

arXiv:2606.01436v1 Announce Type: new Abstract: The growing capabilities of large language models (LLMs) have led to the saturation of many benchmarks and training datasets used to improve them. Motiv

researcharxiv-cs-cl
2 Jun 2026
Agents

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

DGX agent

arXiv:2602.03619v2 Announce Type: replace Abstract: Nowadays, developing reliable DeepResearch-style long-form report generation remains challenging, as training and evaluation lack verifiable reward

agentsarxiv-cs-cl
2 Jun 2026
Model Releases

Learning to Retrieve: Dual-Level Long-Term Memory for Text-to-SQL Agents

DGX agent

arXiv:2606.00547v1 Announce Type: new Abstract: Interactive text-to-SQL agents solve database tasks through multi-turn interactions involving schema exploration, query execution, feedback interpretati

model-releasesarxiv-cs-cl
2 Jun 2026
Research

Lessons from the Trenches on Reproducible Evaluation of Language Models

DGX agent

arXiv:2405.14782v3 Announce Type: replace Abstract: Reliable evaluation of language models (LMs) remains an open challenge. Re- searchers and engineers face methodological issues such as the sensitivi

researcharxiv-cs-cl
2 Jun 2026
Research

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

DGX agent

arXiv:2602.23881v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to propose candidate t

researcharxiv-cs-cl
2 Jun 2026
Safety

LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation

DGX agent

arXiv:2603.09403v2 Announce Type: replace Abstract: Validating evaluation metrics for NLG typically relies on expensive and time-consuming human annotations, which predominantly exist only for English

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection

DGX agent

arXiv:2606.00684v1 Announce Type: cross Abstract: We address the problem of out-of-distribution (OOD) detection for target observations embedded in a subspace of the high dimensional data space. Using

model-releasesarxiv-cs-cl
2 Jun 2026
Applications

LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning

DGX agent

arXiv:2606.01336v1 Announce Type: new Abstract: As real-world applications increasingly require processing inputs of 100k+ tokens, the gap between context length and inference efficiency has become a

applicationsarxiv-cs-cl
2 Jun 2026
Safety

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

DGX agent

arXiv:2606.00975v1 Announce Type: new Abstract: LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delu

safetyarxiv-cs-cl
2 Jun 2026
Research

M^3 Scaling Law: Optimizing Multi-Epoch, Multi-Lingual, and Multi-Stage Training for Low-Resource Language Models

DGX agent

arXiv:2410.12325v2 Announce Type: replace Abstract: In this paper, we study a fundamental design problem in pretraining Large Language Models (LLMs) for low-resource language regimes. Existing works a

researcharxiv-cs-cl
2 Jun 2026
Research

Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling

DGX agent

arXiv:2606.02004v1 Announce Type: new Abstract: Consumer-price measurement increasingly draws on alternative data sources -- scanner, web-scraped, and transaction/receipt data. A recurring obstacle is

researcharxiv-cs-cl
2 Jun 2026
Research

Malaysian English News Decoded: A Linguistic Resource for Named Entity and Relation Extraction

DGX agent

arXiv:2402.14521v2 Announce Type: replace Abstract: Standard English and Malaysian English exhibit notable differences, posing challenges for natural language processing (NLP) tasks on Malaysian Engli

researcharxiv-cs-cl
2 Jun 2026
Model Releases

MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation

DGX agent

arXiv:2505.18614v5 Announce Type: replace Abstract: Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style. In animated mu

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning

DGX agent

arXiv:2606.01914v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) remain unreliable on spatial multiple-choice questions, and their failures are often attributed to poorly atten

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning

DGX agent

arXiv:2606.01301v1 Announce Type: new Abstract: Hallucinations in medical large language models (LLMs) pose serious risks for clinical decision support, particularly when models must reason over compl

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

MemoNoveltyAgent: A Historical Research Memory-Aware Agent Workflow for Paper Novelty Assessment

DGX agent

arXiv:2603.20884v2 Announce Type: replace Abstract: To alleviate the heavy burden of paper screening, researchers increasingly rely on existing AI agents, such as AI reviewers or DeepResearch, for pap

model-releasesarxiv-cs-cl
2 Jun 2026
Agents

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

DGX agent

arXiv:2602.03318v3 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling-a slow and fragile process ill-suited to novel scenarios. While large language models (LLM

agentsarxiv-cs-cl
2 Jun 2026
Safety

Mitigating Bias in Locally Constrained Decoding via Tractable Proposals

DGX agent

arXiv:2606.01926v1 Announce Type: new Abstract: Generations from large language models often fail to conform to desired constraints such as JSON schema. Existing locally constrained decoding (LCD) app

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

Model-Based Quality Assessment for Massively Multilingual Parallel Data

DGX agent

arXiv:2606.00285v1 Announce Type: new Abstract: Large-scale multilingual bitext often contains two distinct problems: non-parallel sentence pairs and low-quality translations. We decompose model-based

model-releasesarxiv-cs-cl
2 Jun 2026
← Previous
1…6263646566…162
Next →