AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Research

C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

DGX agent

arXiv:2604.13618v1 Announce Type: new Abstract: Rubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification. H

researcharxiv-cs-cl
16 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference

DGX agent

arXiv:2604.13634v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by letting draft tokens bypass full verification, but conventional frameworks suffer from fre

researcharxiv-cs-cl
16 Apr 2026
Model Releases

Can Large Language Models Reliably Extract Physiology Index Values from Coronary Angiography Reports?

DGX agent

arXiv:2604.13077v1 Announce Type: new Abstract: Coronary angiography (CAG) reports contain clinically relevant physiological measurements, yet this information is typically in the form of unstructured

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

CANVAS: Continuity-Aware Narratives via Visual Agentic Storyboarding

DGX agent

arXiv:2604.13452v1 Announce Type: new Abstract: Long-form visual storytelling requires maintaining continuity across shots, including consistent characters, stable environments, and smooth scene trans

model-releasesarxiv-cs-cl
16 Apr 2026
Research

Caption First, VQA Second: Knowledge Density, Not Task Format, Drives Multimodal Scaling

DGX agent

arXiv:2604.13054v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved rapid progress, yet their scaling behavior remains less clearly characterized and often less pred

researcharxiv-cs-cl
16 Apr 2026
Research

Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs

DGX agent

arXiv:2604.13950v1 Announce Type: new Abstract: We show how causal interventions in Transformer models provide insights into English syntax by focusing on a long-standing challenge for syntactic theor

researcharxiv-cs-cl
16 Apr 2026
Model Releases

Chain of Uncertain Rewards with Large Language Models for Reinforcement Learning

DGX agent

arXiv:2604.13504v1 Announce Type: cross Abstract: Designing effective reward functions is a cornerstone of reinforcement learning (RL), yet it remains a challenging and labor-intensive process due to

model-releasesarxiv-cs-cl
16 Apr 2026
Safety

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

DGX agent

arXiv:2603.27064v2 Announce Type: replace-cross Abstract: Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a ca

safetyarxiv-cs-cl
16 Apr 2026
Agents

Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models

DGX agent

arXiv:2604.13706v1 Announce Type: new Abstract: Professional fact-checkers rely on domain knowledge and deep contextual understanding to verify claims. Large language models (LLMs) and large reasoning

agentsarxiv-cs-cl
16 Apr 2026
Model Releases

CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation

DGX agent

arXiv:2504.21751v4 Announce Type: replace-cross Abstract: Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Coherence in the brain unfolds across separable temporal regimes

DGX agent

arXiv:2512.20481v4 Announce Type: replace-cross Abstract: To maintain coherence in language, the brain must satisfy key competing temporal demands: the gradual accumulation of meaning across extended

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

CollabCoder: Plan-Code Co-Evolution via Collaborative Decision-Making for Efficient Code Generation

DGX agent

arXiv:2604.13946v1 Announce Type: cross Abstract: Automated code generation remains a persistent challenge in software engineering, as conventional multi-agent frameworks are often constrained by stat

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Common to Whom? Regional Cultural Commonsense and LLM Bias in India

DGX agent

arXiv:2601.15550v3 Announce Type: replace Abstract: Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commo

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Correct Chains, Wrong Answers: Dissociating Reasoning from Output in LLM Logic

DGX agent

arXiv:2604.13065v1 Announce Type: new Abstract: LLMs can execute every step of chain-of-thought reasoning correctly and still produce wrong final answers. We introduce the Novel Operator Test, a bench

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

DGX agent

arXiv:2604.14121v1 Announce Type: new Abstract: LLM reasoning traces suffer from complex flaws -- *Step Internal Flaws* (logical errors, hallucinations, etc.) and *Step-wise Flaws* (overthinking, unde

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

DGX agent

arXiv:2405.19088v3 Announce Type: replace Abstract: Recent advancements in large multimodal language models have demonstrated remarkable proficiency across a wide range of tasks. Yet, these models sti

model-releasesarxiv-cs-cl
16 Apr 2026
Research

Curation of a Palaeohispanic Dataset for Machine Learning

DGX agent

arXiv:2604.13070v1 Announce Type: new Abstract: Palaeohispanic languages are those spoken in the Iberian Peninsula before the arrival of the Romans in the 3rd Century B.C. Their study was really put o

researcharxiv-cs-cl
16 Apr 2026
Safety

Debate to Align: Reliable Entity Alignment through Two-Stage Multi-Agent Debate

DGX agent

arXiv:2604.13551v1 Announce Type: new Abstract: Entity alignment (EA) aims to identify entities referring to the same real-world object across different knowledge graphs (KGs). Recent approaches based

safetyarxiv-cs-cl
16 Apr 2026
Research

Deep Learning Based Amharic Chatbot for FAQs in Universities

DGX agent

arXiv:2402.01720v3 Announce Type: replace-cross Abstract: University students often spend a considerable amount of time seeking answers to common questions from administrators or teachers. This can be

researcharxiv-cs-cl
16 Apr 2026
Model Releases

DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs

DGX agent

arXiv:2604.13075v1 Announce Type: new Abstract: Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

DGX agent

arXiv:2604.13060v1 Announce Type: new Abstract: Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiogr

model-releasesarxiv-cs-cl
16 Apr 2026
Tutorials

Diffusion Language Models for Speech Recognition

DGX agent

arXiv:2604.14001v1 Announce Type: new Abstract: Diffusion language models have recently emerged as a leading alternative to standard language models, due to their ability for bidirectional attention a

tutorialsarxiv-cs-cl
16 Apr 2026
Model Releases

Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection

DGX agent

arXiv:2604.13899v1 Announce Type: new Abstract: Instruction-tuned LLMs can annotate thousands of instances from a short prompt at negligible cost. This raises two questions for active learning (AL): c

model-releasesarxiv-cs-cl
16 Apr 2026
Safety

Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA

DGX agent

arXiv:2604.13731v1 Announce Type: new Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existin

safetyarxiv-cs-cl
16 Apr 2026
Model Releases

Document-tuning for robust alignment to animals

DGX agent

arXiv:2604.13076v1 Announce Type: new Abstract: We investigate the robustness of value alignment via finetuning with synthetic documents, using animal compassion as a value that is both important in i

model-releasesarxiv-cs-cl
16 Apr 2026
Research

Dual-Enhancement Product Bundling: Bridging Interactive Graph and Large Language Model

DGX agent

arXiv:2604.14030v1 Announce Type: new Abstract: Product bundling boosts e-commerce revenue by recommending complementary item combinations. However, existing methods face two critical challenges: (1)

researcharxiv-cs-cl
16 Apr 2026
Research

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints

DGX agent

arXiv:2604.13371v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly described as possessing strong reasoning capabilities, supported by high performance on mathematical, logi

researcharxiv-cs-cl
16 Apr 2026
Research

English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training

DGX agent

arXiv:2604.13286v1 Announce Type: new Abstract: Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to p

researcharxiv-cs-cl
16 Apr 2026
Model Releases

Evaluating LLM-Based Translation of a Low-Resource Technical Language: The Medical and Philosophical Greek of Galen

DGX agent

arXiv:2602.24119v2 Announce Type: replace Abstract: Purpose: This study evaluates the quality of commercial large language model (LLM) machine translation (MT) for Ancient Greek technical prose and be

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection

DGX agent

arXiv:2604.13232v1 Announce Type: new Abstract: This discussion paper re-examines SemEval-2020 Task 1, the most influential shared benchmark for lexical semantic change detection, through a three-part

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy

DGX agent

arXiv:2604.02709v2 Announce Type: replace Abstract: The formal reasoning capabilities of LLMs are crucial for advancing automated software engineering. However, existing benchmarks for LLMs lack syste

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

EVE: A Domain-Specific LLM Framework for Earth Intelligence

DGX agent

arXiv:2604.13071v1 Announce Type: new Abstract: We introduce Earth Virtual Expert (EVE), the first open-source, end-to-end initiative for developing and deploying domain-specialized LLMs for Earth Int

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Exposia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

DGX agent

arXiv:2601.06536v2 Announce Type: replace Abstract: We present Exposia, the first public dataset that connects writing and feedback in higher education, enabling research on educationally grounded com

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

ExpSeek: Self-Triggered Experience Seeking for Web Agents

DGX agent

arXiv:2601.08605v2 Announce Type: replace Abstract: Experience intervention in web agents emerges as a promising technical paradigm, enhancing agent interaction capabilities by providing valuable insi

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

F-Actor: Controllable Conversational Behaviour in Full-Duplex Models

DGX agent

arXiv:2601.11329v3 Announce Type: replace Abstract: Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions

DGX agent

arXiv:2509.18847v3 Announce Type: replace-cross Abstract: Tool-augmented large language models (LLMs) are usually trained with supervised imitation or coarse-grained reinforcement learning that optimi

model-releasesarxiv-cs-cl
16 Apr 2026
Local Ai

Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments

DGX agent

arXiv:2508.08791v3 Announce Type: replace Abstract: Effective tool use is essential for large language models (LLMs) to interact with their environment. However, progress is limited by the lack of eff

local-aiarxiv-cs-cl
16 Apr 2026
Model Releases

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

DGX agent

arXiv:2511.21760v3 Announce Type: replace Abstract: Recent advances in multimodal large language models (LLMs) have enabled unified reasoning across images, audio, and video, but extending such capabi

model-releasesarxiv-cs-cl
16 Apr 2026
Safety

Foresight Optimization for Strategic Reasoning in Large Language Models

DGX agent

arXiv:2604.13592v1 Announce Type: new Abstract: Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-ba

safetyarxiv-cs-cl
16 Apr 2026
Agents

Form Without Function: Agent Social Behavior in the Moltbook Network

DGX agent

arXiv:2604.13052v1 Announce Type: cross Abstract: Moltbook is a social network where every participant is an AI agent. We analyze 1,312,238 posts, 6.7~million comments, and over 120,000 agent profiles

agentsarxiv-cs-cl
16 Apr 2026
Local Ai

From Anchors to Supervision: Memory-Graph Guided Corpus-Free Unlearning for Large Language Models

DGX agent

arXiv:2604.13777v1 Announce Type: new Abstract: Large language models (LLMs) may memorize sensitive or copyrighted content, raising significant privacy and legal concerns. While machine unlearning has

local-aiarxiv-cs-cl
16 Apr 2026
Model Releases

From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

DGX agent

arXiv:2604.14137v1 Announce Type: new Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vibe-testing'':

model-releasesarxiv-cs-cl
16 Apr 2026
Research

From Prediction to Justification: Aligning Sentiment Reasoning with Human Rationale via Reinforcement Learning

DGX agent

arXiv:2604.13398v1 Announce Type: new Abstract: While Aspect-based Sentiment Analysis (ABSA) systems have achieved high accuracy in identifying sentiment polarities, they often operate as 'black boxes

researcharxiv-cs-cl
16 Apr 2026
Safety

From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space

DGX agent

arXiv:2604.14142v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), it

safetyarxiv-cs-cl
16 Apr 2026
Applications

From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines

DGX agent

arXiv:2604.13468v1 Announce Type: cross Abstract: Generative information retrieval (GenIR) formulates the retrieval process as a text-to-text generation task, leveraging the vast knowledge of large la

applicationsarxiv-cs-cl
16 Apr 2026
Safety

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

DGX agent

arXiv:2604.13067v1 Announce Type: cross Abstract: SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations ofte

safetyarxiv-cs-cl
16 Apr 2026
Model Releases

From Weights to Activations: Is Steering the Next Frontier of Adaptation?

DGX agent

arXiv:2604.14090v1 Announce Type: new Abstract: Post-training adaptation of language models is commonly achieved through parameter updates or input-based methods such as fine-tuning, parameter-efficie

model-releasesarxiv-cs-cl
16 Apr 2026
Safety

From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution

DGX agent

arXiv:2604.14053v1 Announce Type: new Abstract: Efficiency and safety of Large Language Models (LLMs), among other factors, rely on the quality of tokenization. A good tokenizer not only improves infe

safetyarxiv-cs-cl
16 Apr 2026
← Previous
1…145146147148149…160
Next →