AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
11 May 2026

SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following

Model ReleasesDGX agent

arXiv:2605.06353v2 Announce Type: replace Abstract: In a conversation, a helpful assistant must reliably follow user directives, even as they refine, modify, or contradict earlier requests. Yet most i

Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise

ResearchDGX agent

arXiv:2602.07425v2 Announce Type: replace-cross Abstract: While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have re

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation

SafetyDGX agent

arXiv:2605.07711v1 Announce Type: new Abstract: On-policy distillation (OPD) is a standard tool for transferring teacher behavior to a smaller student, but it implicitly assumes that teacher and stude


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SmellBench: Evaluating LLM Agents on Architectural Code Smell Repair

Model ReleasesDGX agent

arXiv:2605.07001v1 Announce Type: cross Abstract: Architectural code smells erode software maintainability and are costly to repair manually, yet unlike localized bugs, they require cross-module reaso

Sparser, Faster, Lighter Transformer Language Models

HardwareDGX agent

arXiv:2603.23198v2 Announce Type: replace-cross Abstract: Scaling autoregressive large language models (LLMs) has driven unprecedented progress but comes with vast computational costs. In this work, w

SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting

ResearchDGX agent

arXiv:2605.07243v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting a tree of candidate continuations and verifying it in one target forward. Existing drafters f

SSP-based construction of evaluation-annotated data for fine-grained aspect-based sentiment analysis

ResearchDGX agent

arXiv:2605.07446v1 Announce Type: new Abstract: We report the construction of a Korean evaluation-annotated corpus, hereafter called 'Evaluation Annotated Dataset (EVAD)', and its use in Aspect-Based

Statistical Patterns in the Equations of Physics and the Emergence of a Meta-Law of Nature

ResearchDGX agent

arXiv:2408.11065v2 Announce Type: replace-cross Abstract: Physics seeks to uncover the laws of Nature and express them through mathematical equations. Despite the vast diversity of natural phenomena,

TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP

Model ReleasesDGX agent

arXiv:2605.06886v1 Announce Type: new Abstract: This work introduces TajPersLexon, a curated Tajik--Persian parallel lexical resource of 40,112 word and short-phrase pairs for cross-script lexical ret

TCMIIES: A Browser-Based LLM-Powered Intelligent Information Extraction System for Academic Literature

Local AiDGX agent

arXiv:2605.07507v1 Announce Type: new Abstract: The exponential growth of academic publications has created an urgent need for automated tools capable of extracting structured knowledge from unstructu

Teaching Language Models to Think in Code

Model ReleasesDGX agent

arXiv:2605.07237v1 Announce Type: new Abstract: Tool-integrated reasoning (TIR) has emerged as a dominant paradigm for mathematical problem solving in language models, combining natural language (NL)

TextLDM: Language Modeling with Continuous Latent Diffusion

SafetyDGX agent

arXiv:2605.07748v1 Announce Type: new Abstract: Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next st

The Position Curse: LLMs Struggle to Locate the Last Few Items in a List

Model ReleasesDGX agent

arXiv:2605.07127v1 Announce Type: cross Abstract: Modern large language models (LLMs) can find a needle in a haystack (locating a single relevant fact buried among hundreds of thousands of irrelevant

The Proxy Presumption: From Semantic Embeddings to Valid Social Measures

SafetyDGX agent

arXiv:2605.07409v1 Announce Type: new Abstract: Natural Language Processing is rapidly evolving into a primary instrument for Computational Social Science, with researchers increasingly using embeddin

Theoretical Limits of Language Model Alignment

SafetyDGX agent

arXiv:2605.07105v1 Announce Type: cross Abstract: Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance

SafetyDGX agent

arXiv:2605.07461v1 Announce Type: new Abstract: Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for re

Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings

Model ReleasesDGX agent

arXiv:2605.07158v1 Announce Type: cross Abstract: Vector search and retrieval-augmented generation (RAG) rest on the assumption that cosine similarity between text embeddings reflects conceptual relat

Topology-Enhanced Alignment for Large Language Models: Trajectory Topology Loss and Topological Preference Optimization

Local AiDGX agent

arXiv:2605.07172v1 Announce Type: new Abstract: Alignment of large language models (LLMs) via SFT and RLHF/DPO typically ignores the global geometry of the representation space, relying instead on loc

Towards Closing the Autoregressive Gap in Language Modeling via Entropy-Gated Continuous Bitstream Diffusion

Model ReleasesDGX agent

arXiv:2605.07013v1 Announce Type: new Abstract: Diffusion language models (DLMs) promise parallel, order-agnostic generation, but on standard benchmarks they have historically lagged behind autoregres

Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs

ResearchDGX agent

arXiv:2605.07568v1 Announce Type: cross Abstract: The Arrow-of-Time (AoT) task, determining whether a video plays forward or backward by recognizing temporal irreversibility, is one humans solve with

Training-Free Multimodal Large Language Model Orchestration

SafetyDGX agent

arXiv:2508.10016v3 Announce Type: replace Abstract: Building interactive omni-modal assistants often relies on end-to-end multimodal alignment to fuse heterogeneous modalities, which incurs substantia

UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function

SafetyDGX agent

arXiv:2410.21438v3 Announce Type: replace Abstract: By pretraining on trillions of tokens, an LLM gains the capability of text generation. However, to enhance its utility and reduce potential harm, SF

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types

SafetyDGX agent

arXiv:2408.15339v4 Announce Type: replace-cross Abstract: RL alignment methods, including RLHF and DPO, are primarily based on pairwise preference data. Although scalar or score-based feedback has bee

Uncertainty-Aware Structured Data Extraction from Full CMR Reports via Distilled LLMs

ResearchDGX agent

arXiv:2605.08045v1 Announce Type: new Abstract: Converting free-text cardiac magnetic resonance (CMR) reports into auditable structured data remains a bottleneck for cohort assembly, longitudinal cura

User eXperience Perception Insights Dataset (UXPID): Synthetic User Feedback from Public Industrial Forums

ApplicationsDGX agent

arXiv:2509.11777v2 Announce Type: replace Abstract: Customer feedback in industrial forums offers rich but underexplored insights into real-world product experience. Yet systematic analysis remains ch

Utility-Preserving De-Identification for Math Tutoring: Investigating Numeric Ambiguity in the MathEd-PII Benchmark Dataset

Model ReleasesDGX agent

arXiv:2602.16571v2 Announce Type: replace Abstract: Large-scale sharing of dialogue data is key to advancing the science of teaching and learning, yet rigorous de-identification remains a major barrie

WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation

ResearchDGX agent

arXiv:2605.07522v1 Announce Type: new Abstract: Accurate weather forecast reporting enables individuals and communities to better plan daily activities and agricultural operations. However, the curren

When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models

Model ReleasesDGX agent

arXiv:2605.07260v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models route each token to a small subset of experts, but whether the routes selected by a trained top-k router are

When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents

Model ReleasesDGX agent

arXiv:2605.06731v1 Announce Type: cross Abstract: Personalized LLM agents maintain persistent cross-session state to support long-horizon collaboration. Yet, this persistence introduces a subtle but c

Why do Large Language Models Fail in Low-resource Translation? Unraveling the Token Dynamics of Large Language Models for Machine Translation

ResearchDGX agent

arXiv:2605.07533v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently demonstrated strong performance in machine translation (MT). However, most prior work focuses on improving or

WorldCup Sampling for Multi-bit LLM Watermarking

ResearchDGX agent

arXiv:2602.01752v2 Announce Type: replace Abstract: As large language models (LLMs) generate increasingly human-like text, watermarking has emerged as a promising solution for reliable attribution bey

7 May 2026

A Comparative Analysis of Machine Learning and Deep Learning Models for Tweet Sentiment Classification: A Case Study on the Sentiment140 Dataset

ApplicationsDGX agent

arXiv:2605.04888v1 Announce Type: new Abstract: The exponential growth of social media has created an urgent need for automated systems to analyze unstructured public sentiment in real time. This stud

A Comparative Study of PyCaret AutoML and CNN-BiLSTM for Binary Hate Speech Detection in Indonesian Twitter

Model ReleasesDGX agent

arXiv:2605.04885v1 Announce Type: new Abstract: This paper compares a PyCaret AutoML branch and a CNN-BiLSTM branch for binary hate speech detection on Indonesian Twitter using the HS label from the c

A Hybrid Method for Low-Resource Named Entity Recognition

ApplicationsDGX agent

arXiv:2605.04489v1 Announce Type: cross Abstract: Named Entity Recognition (NER) is a critical component of Natural Language Processing with diverse applications in information extraction and conversa

Adapt to Thrive! Adaptive Power-Mean Policy Optimization for Improved LLM Reasoning

SafetyDGX agent

arXiv:2605.04066v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an essential paradigm that enhances the reasoning capabilities of Large Language Models (LLMs).

Adapting Large Language Models to a Low-Resource Agglutinative Language: A Comparative Study of LoRA and QLoRA for Bashkir

Model ReleasesDGX agent

arXiv:2605.04948v1 Announce Type: new Abstract: This paper presents a comparative study of parameter-efficient fine-tuning (PEFT) methods, including LoRA and QLoRA, applied to the task of adapting lar

Anticipating Innovation Using Large Language Models

SafetyDGX agent

arXiv:2605.04875v1 Announce Type: new Abstract: Forecasting innovation, intended as the emergence of new technological combinations, is a fundamental challenge for science and policy. We show that for

Are LLMs Ready for Conflict Monitoring? Empirical Evidence from West Africa

Model ReleasesDGX agent

arXiv:2605.04177v1 Announce Type: new Abstract: As LLMs enter conflict monitoring, understanding systematic distortions in their outputs is critical for humanitarian accountability. We evaluate four v

Assessing Cognitive Effort in L2 Idiomatic Processing: An Eye-Tracking Dataset

Model ReleasesDGX agent

arXiv:2605.04857v1 Announce Type: new Abstract: This paper presents the development and validation of an eye-tracking dataset designed to investigate how second-language (L2) learners process idiomati

Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models

ApplicationsDGX agent

arXiv:2605.05090v1 Announce Type: new Abstract: We present an automated, contrastive evaluation pipeline for auditing the behavioral impact of interventions on large language models. Given a base mode

Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

ResearchDGX agent

arXiv:2605.04128v1 Announce Type: cross Abstract: We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing

Balanced Aggregation: Understanding and Fixing Aggregation Bias in GRPO

SafetyDGX agent

arXiv:2605.04077v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a central paradigm for improving reasoning and code generation in large language mode

Benchmarking POS Tagging for the Tajik Language: A Comparative Study of Neural Architectures on the TajPersParallel Corpus

Model ReleasesDGX agent

arXiv:2605.04576v1 Announce Type: new Abstract: This paper presents the first benchmark for the task of automatic part-of-speech (POS) tagging for the Tajik language. Despite the existence of multilin

BenCSSmark: Making the Social Sciences Count in LLM Research

Model ReleasesDGX agent

arXiv:2605.04886v1 Announce Type: new Abstract: This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation a

Beyond Exponential Decay: Rethinking Error Accumulation in Large Language Models

ResearchDGX agent

arXiv:2505.24187v2 Announce Type: replace Abstract: The prevailing assumption of an exponential decay in large language model (LLM) reliability with sequence length, predicated on independent per-toke

Beyond Public Access in LLM Pre-Training Data

SafetyDGX agent

arXiv:2505.00020v2 Announce Type: replace Abstract: Using a legally obtained dataset of 34 copyrighted O'Reilly Media books, we apply the DE-COP membership inference attack method to investigate wheth

Beyond Semantics: An Evidential Reasoning-Aware Multi-View Learning Framework for Trustworthy Mental Health Prediction

ApplicationsDGX agent

arXiv:2605.05121v1 Announce Type: new Abstract: Automated mental health prediction using textual data has shown promising results with deep learning and large language models. However, deploying these

Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference

Model ReleasesDGX agent

arXiv:2605.04341v1 Announce Type: cross Abstract: We study distillation for large language models under explicit compute constraints, with the goal of producing student models that are not only cheape

CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2605.04495v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) depends on document ranking to provide useful evidence for generation, but conventional reranking methods mainly op

CHE-TKG: Collaborative Historical Evidence and Evolutionary Dynamics Learning for Temporal Knowledge Graph Reasoning

SafetyDGX agent

arXiv:2605.04652v1 Announce Type: new Abstract: Temporal knowledge graph (TKG) reasoning aims to predict future events from historical facts. A key challenge lies in jointly capturing two sources of p

Conceptors for Semantic Steering

Model ReleasesDGX agent

arXiv:2605.04980v1 Announce Type: cross Abstract: Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction who

Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors

Model ReleasesDGX agent

arXiv:2512.06393v5 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve high accuracy on many reasoning benchmarks but remain brittle under structural perturbations of rule-base

Connecting online criminal behavior with machine learning: Using authorship attribution to analyze and link potential online traffickers

SafetyDGX agent

arXiv:2605.04080v1 Announce Type: new Abstract: This research investigated how online criminal activities can be better understood and connected using data-driven machine learning methods. Many illega

Continual Knowledge Updating in LLM Systems: Learning Through Multi-Timescale Memory Dynamics

ResearchDGX agent

arXiv:2605.05097v1 Announce Type: cross Abstract: LLMs are trained once, then deployed into a world that never stops changing. External memory compensates for this, but most systems manage it explicit

Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUs

HardwareDGX agent

arXiv:2605.04357v1 Announce Type: cross Abstract: The usage of large language models (LLMs) has grown increasingly fragmented, with no single model dominating. Meanwhile, cloud providers offer a wide

Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillation

ResearchDGX agent

arXiv:2512.14954v2 Announce Type: replace Abstract: Computing next-token likelihood ratios between two language models (LMs) is a standard task in training paradigms such as knowledge distillation. Si

Detecting Hallucinations in Large Language Models via Internal Attention Divergence Signals

ResearchDGX agent

arXiv:2605.05025v1 Announce Type: new Abstract: We propose a lightweight and single-pass uncertainty quantification method for detecting hallucinations in Large Language Models. The method uses attent

DFPO: Scaling Value Modeling via Distributional Flow towards Robust and Generalizable LLM Post-Training

SafetyDGX agent

arXiv:2602.05890v2 Announce Type: replace-cross Abstract: Training reinforcement learning (RL) systems in real-world environments remains challenging due to noisy supervision and poor out-of-domain (O

DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation

ResearchDGX agent

arXiv:2512.20773v4 Announce Type: replace Abstract: Realistic user simulation is crucial for training and evaluating multi-turn dialogue systems, yet creating simulators that accurately replicate huma

Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs

TutorialsDGX agent

arXiv:2510.09885v5 Announce Type: replace Abstract: Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text o

← Previous
1…8283848586…129
Next →