AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Research

GraphReview: Scientific Paper Evaluation via LLM-Based Graph Message Passing

DGX agent

arXiv:2605.27204v1 Announce Type: new Abstract: Scientific paper evaluation often involves not only assessing a manuscript itself, but also relating it to contemporaneous research and prior literature

researcharxiv-cs-cl
27 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models

DGX agent

arXiv:2605.26575v1 Announce Type: new Abstract: Multilingual embedding models are deployed under the assumption that cross-lingual retrieval is symmetric: if a query in language A retrieves its transl

model-releasesarxiv-cs-cl
27 May 2026
Tutorials

In-Context Optimization for Retrieval-Augmented Generation: A Gradient-Descent Perspective

DGX agent

arXiv:2605.26356v1 Announce Type: new Abstract: In-context learning has recently been linked to implicit gradient descent in linear self-attention models, suggesting that context can induce a forward-

tutorialsarxiv-cs-cl
27 May 2026
Model Releases

InfoSynth: Information-Guided Benchmark Synthesis for LLMs

DGX agent

arXiv:2601.00575v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant advancements in reasoning and code generation, but efficiently creating new benchmarks to

model-releasesarxiv-cs-cl
27 May 2026
Agents

Interactive Agents: Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions

DGX agent

arXiv:2408.15787v2 Announce Type: replace Abstract: Creating effective dialogue systems for mental health support requires high-quality multi-turn counseling dialogue data, yet collecting real counsel

agentsarxiv-cs-cl
27 May 2026
Safety

KARMA: Karma-Aligned Reward Model Adaptation

DGX agent

arXiv:2605.26738v1 Announce Type: new Abstract: Human communication depends on implicit social signals where effectiveness is shaped by tone, context, and conversational norms rather than semantic con

safetyarxiv-cs-cl
27 May 2026
Safety

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

DGX agent

arXiv:2605.26947v1 Announce Type: new Abstract: Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt datase

safetyarxiv-cs-cl
27 May 2026
Model Releases

LaRe: Latent Refocusing for Multimodal Reasoning

DGX agent

arXiv:2511.02360v4 Announce Type: replace-cross Abstract: Chain of Thought (CoT) reasoning enhances logical performance by decomposing complex tasks, yet its multimodal extension faces a trade-off. Th

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Large Language Model-Powered Query-Driven Event Timeline Summarization in Industrial Search

DGX agent

arXiv:2605.27066v1 Announce Type: new Abstract: Understanding how events evolve over time is essential for search engines handling queries about trending news. We present QDET (Query-Driven Event Time

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior

DGX agent

arXiv:2605.26797v1 Announce Type: cross Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden st

model-releasesarxiv-cs-cl
27 May 2026
Research

LATTE: Forecasting Peer Anchored Preference Trajectories for Personalized LLM Generation

DGX agent

arXiv:2605.26612v1 Announce Type: new Abstract: Personalized generation with frozen large language models requires a conditioning signal that is both compact and current. Existing personalization meth

researcharxiv-cs-cl
27 May 2026
Agents

Learning GUI Grounding with Spatial Reasoning from Visual Feedback

DGX agent

arXiv:2509.21552v2 Announce Type: replace-cross Abstract: Graphical User Interface (GUI) grounding is commonly framed as a coordinate prediction task -- given a natural language instruction, generate

agentsarxiv-cs-cl
27 May 2026
Research

Learning to Adapt SFT Data for Better Reasoning Generalization

DGX agent

arXiv:2605.26924v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress, with post-training playing a crucial role in enhancing their reasoning capabilities. Amo

researcharxiv-cs-cl
27 May 2026
Tutorials

Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models

DGX agent

arXiv:2601.03079v4 Announce Type: replace Abstract: Moral sensitivity is the most fundamental capability underlying human moral competence. Although many approaches aim to align large language models

tutorialsarxiv-cs-cl
27 May 2026
Model Releases

Learning to Predict Future-Aligned Research Proposals with Language Models

DGX agent

arXiv:2603.27146v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to assist ideation in research, but evaluating the quality of LLM-generated research proposals re

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

LLMs Are Already Good Tutors: Training-Free Prompt Optimization for Pedagogical Math Tutoring

DGX agent

arXiv:2605.27088v1 Announce Type: new Abstract: Aligning LLMs for math tutoring typically requires RL-based training with multi-GPU infrastructure. We investigate whether training-free prompt optimiza

model-releasesarxiv-cs-cl
27 May 2026
Safety

MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

DGX agent

arXiv:2605.27186v1 Announce Type: new Abstract: Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as the lo

safetyarxiv-cs-cl
27 May 2026
Safety

MATCHA: Matching Text via Contrastive Semantic Alignment

DGX agent

arXiv:2605.27345v1 Announce Type: new Abstract: Reliable evaluation is essential for understanding large language model (LLM) performance, yet today's go-to metrics, namely token-overlap scores (e.g.,

safetyarxiv-cs-cl
27 May 2026
Model Releases

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

DGX agent

arXiv:2601.08267v3 Announce Type: replace Abstract: While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Memory Architectures for Multi-Turn Text-to-SQL: A Benchmark and Empirical Study

DGX agent

arXiv:2605.26394v1 Announce Type: new Abstract: Multi-turn Text-to-SQL is central to enterprise analytics yet remains predominantly evaluated in single-turn settings. We introduce EnterpriseMem-Bench,

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

MerLean-Prover: A Recursive Looping Harness for End-to-End Lean 4 Theorem Proving

DGX agent

arXiv:2605.26959v1 Announce Type: cross Abstract: MerLean-Prover is an end-to-end Lean4 theorem prover that replaces sorry declarations with kernel-checkable proofs. It is built from three agent types

model-releasesarxiv-cs-cl
27 May 2026
Applications

MetaGraph: A Large-Scale Meta-Analysis of GenAI in Financial NLP (2022-2025)

DGX agent

arXiv:2509.09544v3 Announce Type: replace Abstract: Financial NLP has evolved rapidly since late 2022, outpacing narrative surveys. We introduce MetaGraph, a methodology for extracting typed knowledge

applicationsarxiv-cs-cl
27 May 2026
Local Ai

MicroSpec: Accelerating Speculative Decoding with Lightweight In-Context Vocabularies

DGX agent

arXiv:2605.26444v1 Announce Type: new Abstract: Large language models typically employ vocabularies of over 100k tokens, which creates a major computational bottleneck at the final linear projection l

local-aiarxiv-cs-cl
27 May 2026
Tutorials

Model Unlearning Objectives Vary for Distinct Language Functions

DGX agent

arXiv:2605.26454v1 Announce Type: new Abstract: Large language models (LLMs) learn undesirable properties during pretraining, including dangerous knowledge and toxic text generation. Just as post-trai

tutorialsarxiv-cs-cl
27 May 2026
Local Ai

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

DGX agent

arXiv:2605.26842v1 Announce Type: cross Abstract: The Muon optimizer has recently offered a promising alternative to AdamW for large language model training, leveraging matrix orthogonalization to pro

local-aiarxiv-cs-cl
27 May 2026
Model Releases

MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding

DGX agent

arXiv:2605.26320v1 Announce Type: cross Abstract: The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-s

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

NestedKV: Nested Memory Routing for Long-Context KV Cache Compression

DGX agent

arXiv:2605.26678v1 Announce Type: new Abstract: Long-context language models are limited by the memory footprint of the key-value (KV) cache. Existing training-free KV compression methods usually rank

model-releasesarxiv-cs-cl
27 May 2026
Research

Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation

DGX agent

arXiv:2605.27194v1 Announce Type: new Abstract: Distilling demonstration effects into hidden-space interventions offers a lightweight alternative to full finetuning. However, existing multimodal varia

researcharxiv-cs-cl
27 May 2026
Research

NSF-SciFy: Mining the NSF Awards Database for Scientific Claims

DGX agent

arXiv:2503.08600v3 Announce Type: replace Abstract: We introduce NSF-SciFy, a comprehensive dataset of scientific claims and investigation proposals extracted from National Science Foundation award ab

researcharxiv-cs-cl
27 May 2026
Model Releases

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

DGX agent

arXiv:2605.26485v1 Announce Type: cross Abstract: We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-vi

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

On the Hidden Costs of Counterfactual Knowledge Training in LLM Unlearning

DGX agent

arXiv:2605.27083v1 Announce Type: new Abstract: Counterfactual tuning (CFT) has emerged as a promising paradigm for Large Language Model (LLM) unlearning by training models to generate alternative fic

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs

DGX agent

arXiv:2510.05864v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs

model-releasesarxiv-cs-cl
27 May 2026
Tutorials

Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics

DGX agent

arXiv:2605.26840v1 Announce Type: new Abstract: Reinforcement learning with evaluation metrics as rewards is widely used to enhance specific capabilities of language models. However, for tasks such as

tutorialsarxiv-cs-cl
27 May 2026
Model Releases

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

DGX agent

arXiv:2605.26978v1 Announce Type: new Abstract: Text-to-speech (TTS) evaluation for low-resource non-Latin-script languages can fail when it relies on a single ASR round-trip word error rate (WER). A

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark

DGX agent

arXiv:2506.00250v4 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions

DGX agent

arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We in

model-releasesarxiv-cs-cl
27 May 2026
Research

PinPoint: Prompting with Informative Interior Points

DGX agent

arXiv:2605.26689v1 Announce Type: cross Abstract: Modern referring image segmentation pipelines couple a vision-language model (VLM) for grounding with a promptable segmenter such as the Segment Anyth

researcharxiv-cs-cl
27 May 2026
Research

Pop-Up Distractions Reveal Bag-of-Events Behavior in Video Large Language Models

DGX agent

arXiv:2605.27101v1 Announce Type: cross Abstract: A key capability for video understanding is reliably linking subjects to events across time, yet whether Video Large Language Models (VideoLLMs) actua

researcharxiv-cs-cl
27 May 2026
Model Releases

PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers

DGX agent

arXiv:2605.26730v1 Announce Type: new Abstract: The rapid growth in submissions to machine learning venues has strained the scientific peer-review system and intensified interest in LLM-based automate

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Probing Cultural Awareness in LLMs: A Case Study of Cross-Culture Aesthetic Stylistics

DGX agent

arXiv:2605.27296v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in diverse cultural contexts, yet their ability to master aesthetic stylistics, i.e., the strateg

model-releasesarxiv-cs-cl
27 May 2026
Research

Probing Minimalist Phase Structure in LLMs: What Universal Dependencies Cannot Represent

DGX agent

arXiv:2605.26431v1 Announce Type: new Abstract: Structural probes train on Universal Dependencies (UD), which does not encode formal-syntactic abstractions such as phase boundaries or phase-internal c

researcharxiv-cs-cl
27 May 2026
Agents

Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge Extraction

DGX agent

arXiv:2602.00959v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can be seen as compressed knowledge bases, but it remains unclear what knowledge they truly contain and how far t

agentsarxiv-cs-cl
27 May 2026
Model Releases

Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals

DGX agent

arXiv:2605.26999v1 Announce Type: new Abstract: Prompt injection poses a critical threat to the safe deployment of large language models, yet existing detection approaches are typically evaluated unde

model-releasesarxiv-cs-cl
27 May 2026
Research

Psychological Constructs in Shared Semantic Space

DGX agent

arXiv:2605.26801v1 Announce Type: new Abstract: Psychological constructs are often measured in separate instruments, datasets, and research traditions, which makes direct comparison difficult. This pa

researcharxiv-cs-cl
27 May 2026
Research

QAM-W: Joint 2D Codebook Quantization for LLM Weights via Hadamard Rotation and Activation-Aware Scaling

DGX agent

arXiv:2605.26339v1 Announce Type: cross Abstract: Scalar post-training quantizers discard pairwise coordinate structure within weight rows. We introduce QAM-W (Quadrature Amplitude Modulation for Weig

researcharxiv-cs-cl
27 May 2026
Research

Quadratic Term Correction on Heaps' Law

DGX agent

arXiv:2511.14683v2 Announce Type: replace Abstract: Heaps' or Herdan's law characterizes the word-type vs. word-token relation by a power-law function, which is concave in linear-linear scale but a st

researcharxiv-cs-cl
27 May 2026
Research

Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids

DGX agent

arXiv:2605.26770v1 Announce Type: new Abstract: Prior work shows that Large Language Models (LLMs) can transform Explainable AI (XAI) outputs into Natural Language Explanations (NLEs) that score highl

researcharxiv-cs-cl
27 May 2026
Safety

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

DGX agent

arXiv:2605.27315v1 Announce Type: new Abstract: Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language mod

safetyarxiv-cs-cl
27 May 2026
← Previous
1…7475767778…162
Next →