AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Model Releases

FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents

DGX agent

arXiv:2602.01566v2 Announce Type: replace Abstract: Deep research is emerging as a representative long-horizon task for large language model (LLM) agents. However, long trajectories in deep research o

model-releasesarxiv-cs-cl
20 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

GroupDPO: Memory efficient Group-wise Direct Preference Optimization

DGX agent

arXiv:2604.15602v1 Announce Type: new Abstract: Preference optimization is widely used to align Large Language Models (LLMs) with preference feedback. However, most existing methods train on a single

safetyarxiv-cs-cl
20 Apr 2026
Research

How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models

DGX agent

arXiv:2604.15873v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly studied as repositories of linguistic knowledge. In this line of work, models are commonly evaluated both

researcharxiv-cs-cl
20 Apr 2026
Model Releases

HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning

DGX agent

arXiv:2604.15648v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) consistently require new arenas to guide their expanding boundaries, yet their capabilities with hypergraphs remain

model-releasesarxiv-cs-cl
20 Apr 2026
Safety

Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information

DGX agent

arXiv:2604.15701v1 Announce Type: new Abstract: The significant computational demands of large language models have increased interest in distilling reasoning abilities into smaller models via Chain-o

safetyarxiv-cs-cl
20 Apr 2026
Model Releases

IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering

DGX agent

arXiv:2510.23536v2 Announce Type: replace Abstract: Intent identification serves as the foundation for generating appropriate responses in personalized question answering (PQA). However, existing benc

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Is this chart lying to me? Automating the detection of misleading visualizations

DGX agent

arXiv:2508.21675v3 Announce Type: replace Abstract: Misleading visualizations are a potent driver of misinformation on social media and the web. By violating chart design principles, they distort data

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

JFinTEB: Japanese Financial Text Embedding Benchmark

DGX agent

arXiv:2604.15882v1 Announce Type: cross Abstract: We introduce JFinTEB, the first comprehensive benchmark specifically designed for evaluating Japanese financial text embeddings. Existing embedding be

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

LaMSUM: Amplifying Voices Against Harassment through LLM Guided Extractive Summarization of User Incident Reports

DGX agent

arXiv:2406.15809v5 Announce Type: replace Abstract: Citizen reporting platforms help the public and authorities stay informed about sexual harassment incidents. However, the high volume of data shared

model-releasesarxiv-cs-cl
20 Apr 2026
Safety

Language, Place, and Social Media: Geographic Dialect Alignment in New Zealand

DGX agent

arXiv:2604.15744v1 Announce Type: new Abstract: This thesis investigates geographic dialect alignment in place-informed social media communities, focussing on New Zealand-related Reddit communities. B

safetyarxiv-cs-cl
20 Apr 2026
Research

Large Reasoning Models Are (Not Yet) Multilingual Latent Reasoners

DGX agent

arXiv:2601.02996v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve strong performance on mathematical reasoning tasks, often attributed to their capability to generate explicit

researcharxiv-cs-cl
20 Apr 2026
Model Releases

LLMs Corrupt Your Documents When You Delegate

DGX agent

arXiv:2604.15597v1 Announce Type: new Abstract: Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning

DGX agent

arXiv:2604.16058v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Models (LLMs) in software development has made distinguishing AI-generated code from human-written code a cr

model-releasesarxiv-cs-cl
20 Apr 2026
Research

Measuring the Semantic Structure and Evolution of Conspiracy Theories

DGX agent

arXiv:2603.26062v2 Announce Type: replace Abstract: Research on conspiracy theories has largely focused on belief formation, exposure, and diffusion, while paying less attention to how their meanings

researcharxiv-cs-cl
20 Apr 2026
Model Releases

MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents

DGX agent

arXiv:2604.15774v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) with persistent memory enhances interaction continuity and personalization but introduces new safety risks. Speci

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

MUSCAT: MUltilingual, SCientific ConversATion Benchmark

DGX agent

arXiv:2604.15929v1 Announce Type: new Abstract: The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experi

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

No Universal Courtesy: A Cross-Linguistic, Multi-Model Study of Politeness Effects on LLMs Using the PLUM Corpus

DGX agent

arXiv:2604.16275v1 Announce Type: new Abstract: This paper explores the response of Large Language Models (LLMs) to user prompts with different degrees of politeness and impoliteness. The Politeness T

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Olmo Hybrid: From Theory to Practice and Back

DGX agent

arXiv:2604.03444v3 Announce Type: replace-cross Abstract: Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid m

model-releasesarxiv-cs-cl
20 Apr 2026
Safety

On the Rejection Criterion for Proxy-based Test-time Alignment

DGX agent

arXiv:2604.16146v1 Announce Type: new Abstract: Recent works proposed test-time alignment methods that rely on a small aligned model as a proxy that guides the generation of a larger base (unaligned)

safetyarxiv-cs-cl
20 Apr 2026
Model Releases

Optimizing Korean-Centric LLMs via Token Pruning

DGX agent

arXiv:2604.16235v1 Announce Type: new Abstract: This paper presents a systematic benchmark of state-of-the-art multilingual large language models (LLMs) adapted via token pruning - a compression techn

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Predicting Where Steering Vectors Succeed

DGX agent

arXiv:2604.15557v1 Announce Type: cross Abstract: Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Preference Estimation via Opponent Modeling in Multi-Agent Negotiation

DGX agent

arXiv:2604.15687v1 Announce Type: new Abstract: Automated negotiation in complex, multi-party and multi-issue settings critically depends on accurate opponent modeling. However, conventional numerical

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

DGX agent

arXiv:2604.15780v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe b

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Qwen3.5-Omni Technical Report

DGX agent

arXiv:2604.15804v1 Announce Type: new Abstract: In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor,

model-releasesarxiv-cs-cl
20 Apr 2026
Research

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration

DGX agent

arXiv:2604.15945v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to augment the input to Large Language Models (LLMs) with external information, such as recent or do

researcharxiv-cs-cl
20 Apr 2026
Model Releases

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

DGX agent

arXiv:2601.03699v2 Announce Type: replace Abstract: As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount.

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

DGX agent

arXiv:2604.15736v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-maki

model-releasesarxiv-cs-cl
20 Apr 2026
Research

Reward Modeling for Scientific Writing Evaluation

DGX agent

arXiv:2601.11374v2 Announce Type: replace Abstract: Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage

researcharxiv-cs-cl
20 Apr 2026
Model Releases

SCHK-HTC: Sibling Contrastive Learning with Hierarchical Knowledge-Aware Prompt Tuning for Hierarchical Text Classification

DGX agent

arXiv:2604.15998v1 Announce Type: new Abstract: Few-shot Hierarchical Text Classification (few-shot HTC) is a challenging task that involves mapping texts to a predefined tree-structured label hierarc

model-releasesarxiv-cs-cl
20 Apr 2026
Research

Sentiment Analysis of German Sign Language Fairy Tales

DGX agent

arXiv:2604.16138v1 Announce Type: new Abstract: We present a dataset and a model for sentiment analysis of German sign language (DGS) fairy tales. First, we perform sentiment analysis for three levels

researcharxiv-cs-cl
20 Apr 2026
Research

SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space

DGX agent

arXiv:2504.16315v4 Announce Type: replace-cross Abstract: The complexity of Sign Language (SL) data processing brings many challenges. The current approach to recognition of SL signs aims to translate

researcharxiv-cs-cl
20 Apr 2026
Safety

SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding

DGX agent

arXiv:2604.15628v1 Announce Type: cross Abstract: Cross-modal retrieval between food images and recipe texts is an important task with applications in nutritional management, dietary logging, and cook

safetyarxiv-cs-cl
20 Apr 2026
Safety

Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing

DGX agent

arXiv:2604.15771v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a foundational paradigm for grounding large language models in external knowledge. While adaptive re

safetyarxiv-cs-cl
20 Apr 2026
Model Releases

Stochasticity in Tokenisation Improves Robustness

DGX agent

arXiv:2604.16037v1 Announce Type: new Abstract: The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities in perturbations of tokenisation

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

SwanNLP at SemEval-2026 Task 5: An LLM-based Framework for Plausibility Scoring in Narrative Word Sense Disambiguation

DGX agent

arXiv:2604.16262v1 Announce Type: new Abstract: Recent advances in language models have substantially improved Natural Language Understanding (NLU). Although widely used benchmarks suggest that Large

model-releasesarxiv-cs-cl
20 Apr 2026
Research

Target-Oriented Pretraining Data Selection via Neuron-Activated Graph

DGX agent

arXiv:2604.15706v1 Announce Type: new Abstract: Everyday tasks come with a target, and pretraining models around this target is what turns them into experts. In this paper, we study target-oriented la

researcharxiv-cs-cl
20 Apr 2026
Model Releases

The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring

DGX agent

arXiv:2604.15702v1 Announce Type: new Abstract: We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework a

model-releasesarxiv-cs-cl
20 Apr 2026
Research

Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch

DGX agent

arXiv:2604.15490v1 Announce Type: new Abstract: Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks

researcharxiv-cs-cl
20 Apr 2026
Model Releases

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

DGX agent

arXiv:2505.24672v2 Announce Type: replace Abstract: Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploit

model-releasesarxiv-cs-cl
20 Apr 2026
Research

TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models

DGX agent

arXiv:2604.15756v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP exhibit strong Out-of-distribution (OOD) detection capabilities by aligning visual and textual representation

researcharxiv-cs-cl
20 Apr 2026
Research

Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation

DGX agent

arXiv:2511.02626v3 Announce Type: replace Abstract: Prior works have shown that fine-tuning on new knowledge can induce factual hallucinations in large language models (LLMs), leading to incorrect out

researcharxiv-cs-cl
20 Apr 2026
Safety

UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval

DGX agent

arXiv:2604.15827v1 Announce Type: cross Abstract: Conventional information retrieval is concerned with identifying the relevance of texts for a given query. Yet, the conventional definition of relevan

safetyarxiv-cs-cl
20 Apr 2026
Safety

Whose Facts Win? LLM Source Preferences under Knowledge Conflicts

DGX agent

arXiv:2601.03746v3 Announce Type: replace Abstract: As large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their beh

safetyarxiv-cs-cl
20 Apr 2026
Safety

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

DGX agent

arXiv:2408.15549v4 Announce Type: replace Abstract: As large language models (LLMs) continue to advance, aligning these models with human preferences has emerged as a critical challenge. Traditional a

safetyarxiv-cs-cl
20 Apr 2026
Model Releases

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

DGX agent

arXiv:2510.17210v3 Announce Type: replace Abstract: The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Alon

model-releasesarxiv-cs-cl
20 Apr 2026
Research

A Linguistics-Aware LLM Watermarking via Syntactic Predictability

DGX agent

arXiv:2510.13829v3 Announce Type: replace Abstract: As large language models (LLMs) continue to advance rapidly, reliable governance tools have become critical. Publicly verifiable watermarking is par

researcharxiv-cs-cl
17 Apr 2026
Model Releases

AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization

DGX agent

arXiv:2511.15915v2 Announce Type: replace-cross Abstract: We present AccelOpt, a self-improving large language model (LLM) agentic system that autonomously optimizes kernels for emerging AI acclerator

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Acceptance Dynamics Across Cognitive Domains in Speculative Decoding

DGX agent

arXiv:2604.14682v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model (LLM) inference. It uses a small draft model to propose a tree of future tokens. A larger target

model-releasesarxiv-cs-cl
17 Apr 2026
← Previous
1…140141142143144…160
Next →