AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning

DGX agent

arXiv:2510.07731v3 Announce Type: replace-cross Abstract: Organic reaction mechanisms are the stepwise elementary reactions by which reactants form intermediates and products, and are fundamental to u

model-releasesarxiv-cs-cl
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility

DGX agent

arXiv:2605.01357v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at long-context understanding but exhibit significant limitations in long-form generation. Existing studies primarily

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval

DGX agent

arXiv:2506.11499v2 Announce Type: replace Abstract: Multimodal chatbots have become one of the major topics for dialogue systems in both research community and industry. Recently, researchers have she

model-releasesarxiv-cs-cl
5 May 2026
Research

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality

DGX agent

arXiv:2605.01749v1 Announce Type: new Abstract: Large Reasoning Models achieve strong performance on complex tasks but remain prone to hallucinations, particularly in long-form generation where errors

researcharxiv-cs-cl
5 May 2026
Model Releases

OpenAI GPT-5 System Card

DGX agent

arXiv:2601.03267v2 Announce Type: replace Abstract: This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice

DGX agent

arXiv:2605.01333v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-lev

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Orthographic Constraint Satisfaction and Human Difficulty Alignment in Large Language Models

DGX agent

arXiv:2511.21086v2 Announce Type: replace Abstract: Large language models must satisfy hard orthographic constraints during controlled text generation, yet systematic cross-family evaluation remains l

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Pair2Score: Pairwise-to-Absolute Transfer for LLM-Based Essay Scoring

DGX agent

arXiv:2605.02069v1 Announce Type: new Abstract: Many scoring applications require absolute predictions, while pairwise comparisons can provide a simpler learning objective. We present Pair2Score, a tw

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

PC-MNet: Dual-Level Congruity Modeling for Multimodal Sarcasm Detection via Polarity-Modulated Attention

DGX agent

arXiv:2605.02447v1 Announce Type: new Abstract: Multimodal sarcasm detection, which aims to precisely identify pragmatic incongruities between literal text and nonverbal cues, has gained substantial a

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Perturbation Dose Responses in Recursive LLM Loops: Raw Switching, Stochastic Floors, and Persistent Escape under Append, Replace, and Dialog Updates

DGX agent

arXiv:2605.02236v1 Announce Type: cross Abstract: Recursive language-model loops often settle into recognizable attractor-like patterns. The practical question is how much injected text is needed to m

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Prescriptive Scaling Laws for Data Constrained Training

DGX agent

arXiv:2605.01640v1 Announce Type: cross Abstract: Training compute is increasingly outpacing the availability of high-quality data. This shifts the central challenge from optimal compute allocation to

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Pretraining A Large Language Model using Distributed GPUs: A Memory-Efficient Decentralized Paradigm

DGX agent

arXiv:2602.11543v2 Announce Type: replace Abstract: Pretraining large language models (LLMs) typically requires centralized clusters with thousands of high-memory GPUs (e.g., H100/A100). Recent decent

model-releasesarxiv-cs-cl
5 May 2026
Research

Progress Ratio Embeddings: An Impatience Signal for Robust Length Control in Neural Text Generation

DGX agent

arXiv:2512.06938v2 Announce Type: replace Abstract: Modern neural language models achieve high accuracy in text generation, yet precise control over generation length remains underdeveloped. In this p

researcharxiv-cs-cl
5 May 2026
Model Releases

Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese

DGX agent

arXiv:2605.01630v1 Announce Type: new Abstract: Rankings produced by holistic LLM-as-a-judge scoring are sensitive to the bias of the chosen judge model. We show that switching to binary rubric scorin

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Psychologically Potent, Computationally Invisible: LLMs Generate Social-Comparison Triggers They Fail to Detect

DGX agent

arXiv:2605.01017v1 Announce Type: new Abstract: We introduce Xiaohongshu Social Comparison Reader Elicitation (XHS-SCoRE), a reader-grounded benchmark for detecting if a text-only Xiaohongshu (RedNote

model-releasesarxiv-cs-cl
5 May 2026
Research

PubMed-Ophtha: An open resource for training ophthalmology vision-language models on scientific literature

DGX agent

arXiv:2605.02720v1 Announce Type: cross Abstract: Vision-language models hold considerable promise for ophthalmology, but their development depends on large-scale, high-quality image-text datasets tha

researcharxiv-cs-cl
5 May 2026
Research

Quantifying and Predicting Disagreement in Graded Human Ratings

DGX agent

arXiv:2605.01168v1 Announce Type: new Abstract: It is increasingly recognized that human annotators do not always agree, and such disagreement is inherent in many annotation tasks. However, not all in

researcharxiv-cs-cl
5 May 2026
Tutorials

RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions

DGX agent

arXiv:2605.01104v1 Announce Type: cross Abstract: Understanding how developers interact with AI coding assistants requires more than chat logs or git histories in isolation; it requires reconstructing

tutorialsarxiv-cs-cl
5 May 2026
Safety

RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems

DGX agent

arXiv:2509.10746v3 Announce Type: replace Abstract: Large language models in healthcare often produce emotionally flat or opaque responses, failing to provide the transparent reasoning required for cl

safetyarxiv-cs-cl
5 May 2026
Research

ReFRAME or Remain: Unsupervised Lexical Semantic Change Detection with Frame Semantics

DGX agent

arXiv:2602.04514v3 Announce Type: replace Abstract: The majority of contemporary computational methods for lexical semantic change (LSC) detection are based on neural embedding distributional represen

researcharxiv-cs-cl
5 May 2026
Model Releases

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

DGX agent

arXiv:2605.01913v1 Announce Type: cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable t

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces

DGX agent

arXiv:2605.02801v1 Announce Type: new Abstract: As large language model (LLM) agents evolve from isolated tool users into coordinated teams, reinforcement learning (RL) must optimize not only individu

model-releasesarxiv-cs-cl
5 May 2026
Safety

Reinforcement Learning from Compiler and Language Server Feedback

DGX agent

arXiv:2510.22907v2 Announce Type: replace Abstract: Coding agents fail when text-level guesses outrun program facts: they hallucinate APIs, drift to the wrong symbol, and apply edits without evidence

safetyarxiv-cs-cl
5 May 2026
Safety

Reliability-Oriented Multilingual Orthopedic Diagnosis: A Domain-Adaptive Modeling and a Conceptual Validation Framework

DGX agent

arXiv:2605.02266v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly proposed for clinical decision support including multilingual diagnosis in low-resource settings. However,

safetyarxiv-cs-cl
5 May 2026
Applications

ReMedi: Reasoner for Medical Clinical Prediction

DGX agent

arXiv:2605.01474v1 Announce Type: new Abstract: Predicting future clinical outcomes from electronic health records (EHR) remains challenging due to the complexity and heterogeneity of patient data. LL

applicationsarxiv-cs-cl
5 May 2026
Research

Revisiting Semantic Role Labeling: Efficient Structured Inference with Dependency-Informed Analysis

DGX agent

arXiv:2605.02505v1 Announce Type: new Abstract: Semantic Role Labeling (SRL) provides an explicit representation of predicate-argument structure, capturing linguistically grounded relations such as wh

researcharxiv-cs-cl
5 May 2026
Model Releases

RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences

DGX agent

arXiv:2605.01831v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback has become the standard paradigm for language model alignment, where reward models directly determine alignme

model-releasesarxiv-cs-cl
5 May 2026
Research

SCARV: Structure-Constrained Aggregation for Stable Sample Ranking in Redundant NLP Datasets

DGX agent

arXiv:2605.00944v1 Announce Type: cross Abstract: Sample-level rankings are increasingly used in data-centric NLP for analysis, filtering, debugging, and curation, yet existing pipelines typically sco

researcharxiv-cs-cl
5 May 2026
Model Releases

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

DGX agent

arXiv:2605.01489v1 Announce Type: cross Abstract: Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures

DGX agent

arXiv:2605.02601v1 Announce Type: new Abstract: We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an ex

model-releasesarxiv-cs-cl
5 May 2026
Research

Sentiment Analysis of Mobile Legends App Reviews Using Machine Learning and LSTM-Based Deep Learning Models

DGX agent

arXiv:2605.01317v1 Announce Type: new Abstract: This paper compares Machine Learning and LSTM-based Deep Learning methods for sentiment analysis of Mobile Legends app reviews. Using a dataset of 10,00

researcharxiv-cs-cl
5 May 2026
Model Releases

Shadow-Loom: Causal Reasoning over Graphical World Model of Narratives

DGX agent

arXiv:2605.02475v1 Announce Type: cross Abstract: Stories hold a reader's attention because they have causes, secrets, and consequences. Shadow-Loom is an experimental open-source framework that turns

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting

DGX agent

arXiv:2605.02105v1 Announce Type: cross Abstract: Pretraining optimizers are tuned to produce the strongest possible base model, on the assumption that a stronger starting point yields a stronger mode

model-releasesarxiv-cs-cl
5 May 2026
Safety

SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 25+ Sign Languages

DGX agent

arXiv:2605.01720v1 Announce Type: cross Abstract: Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often produced in lab

safetyarxiv-cs-cl
5 May 2026
Agents

Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives

DGX agent

arXiv:2604.06091v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly acting as human delegates in multi-agent environments, where a representative agent integrates di

agentsarxiv-cs-cl
5 May 2026
Research

Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models

DGX agent

arXiv:2605.01853v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate extended solutions, yet it remains unclear whether these traces reflect substantive internal computation or merel

researcharxiv-cs-cl
5 May 2026
Research

SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection

DGX agent

arXiv:2605.02888v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model (LLM) inference by using a small draft model to propose candidate tokens that a larger target mo

researcharxiv-cs-cl
5 May 2026
Model Releases

Spoken Language Identification with Pre-trained Models and Margin Loss

DGX agent

arXiv:2605.01905v1 Announce Type: cross Abstract: For the speaker-controlled spoken language identification task proposed in the TidyLang Challenge 2026, this paper proposes a language identification

model-releasesarxiv-cs-cl
5 May 2026
Safety

SRA: Span Representation Alignment for Large Language Model Distillation

DGX agent

arXiv:2605.01205v1 Announce Type: new Abstract: Cross-Tokenizer Knowledge Distillation (CTKD) enables knowledge transfer between a large language model and a smaller student, even when they employ dif

safetyarxiv-cs-cl
5 May 2026
Model Releases

SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking

DGX agent

arXiv:2605.00974v1 Announce Type: cross Abstract: LLMs are increasingly equipped with safety alignment mechanisms, yet recent studies demonstrate that they remain vulnerable to jailbreaking attacks th

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Storie

DGX agent

arXiv:2601.08510v3 Announce Type: replace Abstract: Movie screenplays are rich long-form narratives that interleave complex character relationships, temporally ordered events, and dialogue-driven inte

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

StressEval: Failure-Driven Dynamic Benchmarking for Knowledge-Intensive Reasoning in Large Language Models

DGX agent

arXiv:2605.01939v1 Announce Type: new Abstract: Static benchmarks for LLMs are increasingly compromised by contamination and overfitting especially on knowledge intensive reasoning tasks While recent

model-releasesarxiv-cs-cl
5 May 2026
Applications

Structural Dilemmas and Developmental Pathways of Legal Argument Mining in the Era of Artificial Intelligence

DGX agent

arXiv:2605.02308v1 Announce Type: new Abstract: Against the backdrop of rapid advances in artificial intelligence, legal argument mining has emerged as an important research area linking legal texts w

applicationsarxiv-cs-cl
5 May 2026
Model Releases

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

DGX agent

arXiv:2508.15658v5 Announce Type: replace Abstract: The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show pr

model-releasesarxiv-cs-cl
5 May 2026
Research

Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations

DGX agent

arXiv:2605.02624v1 Announce Type: new Abstract: There is growing interest in exploring user simulation as an alternative to gathering and scoring real user-chatbot interactions for AI chatbot evaluati

researcharxiv-cs-cl
5 May 2026
Local Ai

TagRAG: Tag-guided Hierarchical Knowledge Graph Retrieval-Augmented Generation

DGX agent

arXiv:2601.05254v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation enhances language models by retrieving external knowledge to support informed and grounded responses. However,

local-aiarxiv-cs-cl
5 May 2026
Model Releases

TCDA: Thread-Constrained Discourse-Aware Modeling for Conversational Sentiment Quadruple Analysis

DGX agent

arXiv:2605.01717v1 Announce Type: new Abstract: Conversational Aspect-based Sentiment Quadruple Analysis (DiaASQ) needs to capture the complex interrelationships in multiple rounds of dialogues. Exist

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines

DGX agent

arXiv:2605.01077v1 Announce Type: new Abstract: Brazil's Unified Health System (SUS) relies on official clinical guidelines that define diagnostic criteria, treatments, dosages, and monitoring procedu

model-releasesarxiv-cs-cl
5 May 2026
← Previous
1…111112113114115…161
Next →