AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing

DGX agent

arXiv:2605.03903v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capabi

model-releasesarxiv-cs-cl
6 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Local Ai

Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards

DGX agent

arXiv:2605.03862v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards has become a common way to improve explicit reasoning in large language models, but final-answer correc

local-aiarxiv-cs-cl
6 May 2026
Model Releases

CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing

DGX agent

arXiv:2605.02910v1 Announce Type: cross Abstract: Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification

DGX agent

arXiv:2605.03476v1 Announce Type: new Abstract: Discharge summaries require extracting critical information from lengthy electronic health records (EHRs), a process that is labor-intensive when perfor

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Detecting Stealth Sycophancy in Mental-Health Dialogue with Dynamic Emotional Signature Graphs

DGX agent

arXiv:2605.03472v1 Announce Type: new Abstract: As conversational AI therapists are increasingly used in psychological support settings, reliable offline evaluation of therapeutic response quality rem

model-releasesarxiv-cs-cl
6 May 2026
Tutorials

Direct Simultaneous Translation Activation for Large Audio-Language Models

DGX agent

arXiv:2509.15692v2 Announce Type: replace-cross Abstract: Simultaneous speech-to-text translation (Simul-S2TT) aims to translate speech into target text in real time, outputting translations while rec

tutorialsarxiv-cs-cl
6 May 2026
Research

Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling

DGX agent

arXiv:2601.21684v2 Announce Type: replace Abstract: Test-Time Scaling enhances the reasoning capabilities of Large Language Models by allocating additional inference compute to broaden the exploration

researcharxiv-cs-cl
6 May 2026
Model Releases

Effective Performance Measurement: Challenges and Opportunities in KPI Extraction from Earnings Calls

DGX agent

arXiv:2605.03147v1 Announce Type: new Abstract: Earnings calls are a key source of financial information about public companies. However, extracting information from these calls is difficult. Unlike t

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage

DGX agent

arXiv:2605.03998v1 Announce Type: new Abstract: Emergency department triage assigns patients an acuity score that determines treatment priority, and clinical evidence documents persistent gender dispa

model-releasesarxiv-cs-cl
6 May 2026
Research

Evaluating Reasoning Models for Queries with Presuppositions

DGX agent

arXiv:2605.03050v1 Announce Type: new Abstract: Millions of users turn to AI models for their information needs. It is conceivable that a large number of user queries contain assumptions that may be f

researcharxiv-cs-cl
6 May 2026
Research

ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability

DGX agent

arXiv:2502.11336v2 Announce Type: replace Abstract: Detecting texts generated by Large Language Models (LLMs) could cause grave mistakes due to incorrect decisions, such as undermining students' acade

researcharxiv-cs-cl
6 May 2026
Model Releases

Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis

DGX agent

arXiv:2605.03441v1 Announce Type: cross Abstract: Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We s

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Feature-Augmented Transformers for Robust AI-Text Detection Across Domains and Generators

DGX agent

arXiv:2605.03969v1 Announce Type: new Abstract: AI-generated text is nowadays produced at scale across domains and heterogeneous generation pipelines, making robustness to distribution shift a central

model-releasesarxiv-cs-cl
6 May 2026
Safety

FINER-SQL: Boosting Small Language Models for Text-to-SQL

DGX agent

arXiv:2605.03465v1 Announce Type: cross Abstract: Large language models have driven major advances in Text-to-SQL generation. However, they suffer from high computational cost, long latency, and data

safetyarxiv-cs-cl
6 May 2026
Research

From prompting to evidence-based translation: A RAG+prompt system for Japanese-Chinese translation and its pedagogical potential

DGX agent

arXiv:2605.03387v1 Announce Type: new Abstract: Large language models perform well on high-resource pairs but are less reliable for Japanese-Chinese sentences containing noun-modifying clause construc

researcharxiv-cs-cl
6 May 2026
Research

Geolocating News about Extreme Climate Events: A Comparative Analysis of Off-the-Shelf Tools for Toponym Identification in German

DGX agent

arXiv:2605.03414v1 Announce Type: new Abstract: Determining the geolocation of extreme climate events and disasters in texts is a common problem in climate impact and adaptation research. Named-entity

researcharxiv-cs-cl
6 May 2026
Model Releases

Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability

DGX agent

arXiv:2605.03196v1 Announce Type: new Abstract: A reliable language model should be able to signal, prior to generation, when a query falls outside its knowledge. We investigate whether representation

model-releasesarxiv-cs-cl
6 May 2026
Applications

GLEAN: Active Generalized Category Discovery with Diverse LLM Feedback

DGX agent

arXiv:2502.18414v2 Announce Type: replace Abstract: Generalized Category Discovery (GCD) is a practical and challenging open-world task that aims to recognize both known and novel categories in unlabe

applicationsarxiv-cs-cl
6 May 2026
Model Releases

Hierarchical Memorization in Large Language Models: Evidence from Citation Generation

DGX agent

arXiv:2511.08877v2 Announce Type: replace Abstract: Large language models (LLMs) generate fluent text across a wide range of tasks, but the fabrication of non-existent academic citations remains a cri

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

How Language Models Process Negation

DGX agent

arXiv:2605.03052v1 Announce Type: new Abstract: We study how Large Language Models (LLMs) process negation mechanistically. First, we establish that even though open-weight models often provide wrong

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic

DGX agent

arXiv:2510.09472v2 Announce Type: replace Abstract: Despite the remarkable progress in neural models, their ability to generalize, a cornerstone for applications such as logical reasoning, remains a c

model-releasesarxiv-cs-cl
6 May 2026
Research

InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacy

DGX agent

arXiv:2507.02974v3 Announce Type: replace-cross Abstract: As major progress in LLM-based long-form text generation enables paradigms such as retrieval-augmented generation (RAG) and inference-time sca

researcharxiv-cs-cl
6 May 2026
Research

Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling

DGX agent

arXiv:2602.00594v2 Announce Type: replace Abstract: A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals t

researcharxiv-cs-cl
6 May 2026
Model Releases

LitVISTA: A Benchmark for Narrative Orchestration in Literary Text

DGX agent

arXiv:2601.06445v2 Announce Type: replace Abstract: Computational narrative analysis aims to capture rhythm, tension, and emotional dynamics in literary texts. Existing large language models can gener

model-releasesarxiv-cs-cl
6 May 2026
Safety

LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models

DGX agent

arXiv:2605.03299v1 Announce Type: new Abstract: Cross-lingual topic modeling aims to discover shared semantic structures across languages, yet existing models depend on sparse bilingual resources and

safetyarxiv-cs-cl
6 May 2026
Applications

Logical Consistency as a Bridge: Improving LLM Hallucination Detection via Label Constraint Modeling between Responses and Self-Judgments

DGX agent

arXiv:2605.03971v1 Announce Type: new Abstract: Large Language Models (LLMs) are prone to factual hallucinations, risking their reliability in real-world applications. Existing hallucination detectors

applicationsarxiv-cs-cl
6 May 2026
Safety

MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

DGX agent

arXiv:2605.03228v1 Announce Type: cross Abstract: As large language model (LLM)-powered agents are increasingly deployed to perform complex, real-world tasks, they face a growing class of attacks that

safetyarxiv-cs-cl
6 May 2026
Model Releases

Maximizing mutual information between prompts and responses improve LLM personalization with no additional data or human oversight

DGX agent

arXiv:2603.19294v2 Announce Type: replace-cross Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labe

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following

DGX agent

arXiv:2605.03858v1 Announce Type: new Abstract: Multi-constraint instruction following requires verifying whether a response satisfies multiple individual requirements, yet LLM judges are often assess

model-releasesarxiv-cs-cl
6 May 2026
Research

Mechanism-Faithful Queueing Simulation Model Translation with Large Language Model Support

DGX agent

arXiv:2601.06543v2 Announce Type: replace Abstract: Queueing simulation studies often require substantial manual effort to translate conceptual system descriptions into executable programs and to veri

researcharxiv-cs-cl
6 May 2026
Model Releases

MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports

DGX agent

arXiv:2605.03103v1 Announce Type: new Abstract: Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudinal medical h

model-releasesarxiv-cs-cl
6 May 2026
Research

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

DGX agent

arXiv:2603.06194v2 Announce Type: replace Abstract: Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-turn inter

researcharxiv-cs-cl
6 May 2026
Hardware

MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier

DGX agent

arXiv:2603.03756v3 Announce Type: replace-cross Abstract: While large language models (LLMs) show promise in scientific discovery, existing research focuses on inference or feedback-driven training, l

hardwarearxiv-cs-cl
6 May 2026
Safety

Multilingual Safety Alignment via Self-Distillation

DGX agent

arXiv:2605.02971v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain hig

safetyarxiv-cs-cl
6 May 2026
Tutorials

Natural Language Processing: A Comprehensive Practical Guide from Tokenisation to RLHF

DGX agent

arXiv:2605.03799v1 Announce Type: new Abstract: This preprint presents a systematic, research-oriented practicum that guides the reader through the entire modern NLP pipeline: from tokenisation and ve

tutorialsarxiv-cs-cl
6 May 2026
Model Releases

Not that Groove: Zero-Shot Symbolic Music Editing

DGX agent

arXiv:2505.08203v2 Announce Type: replace-cross Abstract: While recent advancements in AI music generation have predominantly focused on direct audio synthesis, these systems suffer from inherent rigi

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

OCRR: A Benchmark for Online Correction Recovery under Distribution Shift

DGX agent

arXiv:2605.03153v1 Announce Type: cross Abstract: Static benchmarks measure a model frozen at training time. Real systems face distribution shift: new categories, paraphrased queries, drift: and must

model-releasesarxiv-cs-cl
6 May 2026
Local Ai

On-Device Fine-Tuning via Backprop-Free Zeroth-Order Optimization

DGX agent

arXiv:2511.11362v2 Announce Type: replace-cross Abstract: On-device fine-tuning is a critical capability for edge AI systems, which must support adaptation to different agentic tasks under stringent m

local-aiarxiv-cs-cl
6 May 2026
Model Releases

On Verbalized Confidence Scores for LLMs

DGX agent

arXiv:2412.14737v2 Announce Type: replace Abstract: The rise of large language models (LLMs) and their tight integration into our daily life make it essential to dedicate efforts towards their trustwo

model-releasesarxiv-cs-cl
6 May 2026
Agents

OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories

DGX agent

arXiv:2605.04036v1 Announce Type: cross Abstract: Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominat

agentsarxiv-cs-cl
6 May 2026
Model Releases

PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination

DGX agent

arXiv:2605.03571v1 Announce Type: new Abstract: Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising applicati

model-releasesarxiv-cs-cl
6 May 2026
Research

Permutation-Consensus Listwise Judging for Robust Factuality Evaluation

DGX agent

arXiv:2603.20562v2 Announce Type: replace Abstract: Large language models (LLMs) are now widely used as judges, yet their decisions can change under presentation choices that should be irrelevant. We

researcharxiv-cs-cl
6 May 2026
Model Releases

PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization

DGX agent

arXiv:2605.03129v1 Announce Type: cross Abstract: Browsing-enabled LLM assistants can fetch webpages and answer contact-seeking queries, creating a practical channel for scraping contact-style persona

model-releasesarxiv-cs-cl
6 May 2026
Research

psifx -- Psychological and Social Interactions Feature Extraction Package

DGX agent

arXiv:2407.10266v5 Announce Type: replace Abstract: psifx is a plug-and-play multi-modal feature extraction toolkit, aiming to facilitate and democratize the use of state-of-the-art machine learning t

researcharxiv-cs-cl
6 May 2026
Model Releases

RAG over Thinking Traces Can Improve Reasoning Tasks

DGX agent

arXiv:2605.03344v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) has proven effective for knowledge-intensive tasks, but is widely believed to offer limited benefit for reasoning

model-releasesarxiv-cs-cl
6 May 2026
Applications

Rational Communication Shapes Morphological Composition

DGX agent

arXiv:2605.03510v1 Announce Type: new Abstract: Human languages expand vocabularies by combining existing morphemes rather than inventing arbitrary forms. Communicative efficiency shapes lexical syste

applicationsarxiv-cs-cl
6 May 2026
Model Releases

ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

DGX agent

arXiv:2508.05170v3 Announce Type: replace-cross Abstract: In practice, rigorous reasoning is often a key driver of correct code, while Reinforcement Learning (RL) for code generation often neglects op

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Reproducing Complex Set-Compositional Information Retrieval

DGX agent

arXiv:2605.03824v1 Announce Type: new Abstract: Complex information needs may involve set-compositional queries using conjunction, disjunction, and exclusion, yet it remains unclear whether current re

model-releasesarxiv-cs-cl
6 May 2026
← Previous
1…106107108109110…161
Next →