AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
3,914 results
Hardware

AgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement Learning

DGX agent

arXiv:2606.04484v1 Announce Type: new Abstract: We present AgentJet, a distributed swarm training framework for large language model (LLM) agent reinforcement learning. Unlike centralized frameworks t

hardwarearxiv-cs-ai
4 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms

DGX agent

arXiv:2602.09464v2 Announce Type: replace-cross Abstract: Vericoding refers to the generation of formally verified code from rigorous specifications. Recent AI models show promise in vericoding, but a

model-releasesarxiv-cs-ai
4 Jun 2026
Agents

Archi: Agentic Operations at the CMS Experiment

DGX agent

arXiv:2606.04755v1 Announce Type: cross Abstract: We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of hetero

agentsarxiv-cs-ai
4 Jun 2026
Model Releases

AUDDT: A Unified Benchmark Toolkit for Audio and Speech Deepfake Detectors

DGX agent

arXiv:2509.21597v2 Announce Type: replace-cross Abstract: With the prevalence of artificial intelligence (AI)-generated content, such as audio deepfakes, a large body of recent work has focused on dev

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

Automatic Generation of Titles for Research Papers Using Language Models

DGX agent

arXiv:2606.05085v1 Announce Type: cross Abstract: The title of a research paper conveys its primary idea and, occasionally, its conclusions in a clear and concise manner. Choosing an appropriate title

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?

DGX agent

arXiv:2506.10912v4 Announce Type: replace Abstract: Toxicity remains a leading cause of early-stage drug development failure. Despite advances in molecular design and property prediction, the task of

model-releasesarxiv-cs-ai
4 Jun 2026
Research

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

DGX agent

arXiv:2407.03884v4 Announce Type: replace-cross Abstract: Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and

researcharxiv-cs-ai
4 Jun 2026
Applications

ClustRecNet: A Novel End-to-End Deep Learning Framework for Clustering Algorithm Recommendation

DGX agent

arXiv:2509.25289v4 Announce Type: replace-cross Abstract: Identifying an effective clustering algorithm for a given dataset remains a fundamental unsupervised learning issue. We introduce ClustRecNet,

applicationsarxiv-cs-ai
4 Jun 2026
Model Releases

CodegenBench: Can LLMs Write Efficient Code Across Architectures?

DGX agent

arXiv:2606.04023v1 Announce Type: cross Abstract: While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated enviro

model-releasesarxiv-cs-ai
4 Jun 2026
Research

CounterFace: A Synthetic Face Dataset for Fine-Grained Counterfactual Evaluation of Face Recognition Systems

DGX agent

arXiv:2407.13922v3 Announce Type: replace-cross Abstract: Face recognition (FR) systems are widely deployed in critical applications, making their reliability and robustness across diverse populations

researcharxiv-cs-ai
4 Jun 2026
Model Releases

CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

DGX agent

arXiv:2606.04460v1 Announce Type: cross Abstract: AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabilities. How

model-releasesarxiv-cs-ai
4 Jun 2026
Agents

Description-Code Inconsistency in Real-world MCP Servers: Measurement, Detection, and Security Implications

DGX agent

arXiv:2606.04769v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has emerged as a critical standard empowering Large Language Models (LLMs) to utilize external tools. In this ecosyst

agentsarxiv-cs-ai
4 Jun 2026
Applications

From Motion Signals to Insights: A Unified Framework for Student Behavior Analysis and Feedback in Physical Education Classes

DGX agent

arXiv:2503.06525v2 Announce Type: replace-cross Abstract: Analyzing student behavior in educational scenarios is crucial for enhancing teaching quality and student engagement. Existing AI-based models

applicationsarxiv-cs-ai
4 Jun 2026
Research

GlossAssist -- A Tool to Simplify Corpus Creation and Study the Effect of NLP Models in Low-Resource Documentation Settings

DGX agent

arXiv:2606.04367v1 Announce Type: new Abstract: Interlinear glossed text (IGT) is the standard format for linguistic annotation in language documentation. Producing it manually, however, is often slow

researcharxiv-cs-cl
4 Jun 2026
Research

Handwriting Extraction and Analysis of Signature Lists in Swiss Popular Initiatives

DGX agent

arXiv:2606.05018v1 Announce Type: new Abstract: Popular initiatives and referendums are central to Swiss democracy, yet the validation of handwritten signature lists remains a labor-intensive manual p

researcharxiv-cs-cv
4 Jun 2026
Local Ai

Identifying Gems from Roman RAPIDly

DGX agent

arXiv:2606.05103v1 Announce Type: cross Abstract: The Nancy Grace Roman Space Telescope (Roman), set for launch as early as September 2026, will conduct wide-field infrared imaging surveys with unprec

local-aiarxiv-cs-cv
4 Jun 2026
Agents

MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation

DGX agent

arXiv:2606.04513v1 Announce Type: new Abstract: Lane-level maps are critical infrastructure for autonomous driving and lane-level navigation, yet constructing and maintaining standardized lane network

agentsarxiv-cs-ai
4 Jun 2026
Tutorials

Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

DGX agent

arXiv:2606.04326v1 Announce Type: cross Abstract: Concept bottleneck models predict outcomes from high-level concepts detected in inputs. Although concepts provide a simple way to reap benefits from i

tutorialsarxiv-cs-ai
4 Jun 2026
Model Releases

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

DGX agent

arXiv:2606.04773v1 Announce Type: cross Abstract: Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffe

model-releasesarxiv-cs-cl
4 Jun 2026
Research

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

DGX agent

arXiv:2606.04503v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fu

researcharxiv-cs-ai
4 Jun 2026
Model Releases

StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL Synthesis

DGX agent

arXiv:2606.04246v1 Announce Type: new Abstract: Automatic generation of RTL code for digital hardware designs remains challenging due to long-horizon reasoning, multi-step dependencies, and strict cor

model-releasesarxiv-cs-ai
4 Jun 2026
Applications

StrokeTimer: Robust Representation Learning for Ischemic Stroke Onset-Time Estimation from Non-contrast CT

DGX agent

arXiv:2606.04722v1 Announce Type: new Abstract: Ischemic stroke is a major global disease. Treatment decisions are highly time-sensitive, as eligibility for reperfusion therapies relies on the interva

applicationsarxiv-cs-cv
4 Jun 2026
Safety

VentAgent: When LLMs Learn to Breathe -- Multi-Objective Arbitration for ARDS Ventilation

DGX agent

arXiv:2606.04632v1 Announce Type: cross Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung pr

safetyarxiv-cs-cl
4 Jun 2026
Research

A Hybrid Approach For Malware Classification Using Secondary Features Fusion

DGX agent

arXiv:2606.03432v1 Announce Type: cross Abstract: The number of malware (either variant or novel) is rapidly increasing, making malware detection and mitigation a complex problem. One approach to impr

researcharxiv-cs-ai
3 Jun 2026
Research

AlphaEval: A Comprehensive and Efficient Evaluation Framework for Formula Alpha Mining

DGX agent

arXiv:2508.13174v2 Announce Type: replace Abstract: Formula alpha mining, which generates predictive signals from financial data, is critical for quantitative investment. Although various algorithmic

researcharxiv-cs-ai
3 Jun 2026
Applications

Assessing Pause Thresholds for empirical Translation Process Research

DGX agent

arXiv:2604.01410v2 Announce Type: replace Abstract: Text production (and translations) proceeds in the form of stretches of typing, interrupted by keystroke pauses. It is often assumed that fast typin

applicationsarxiv-cs-cl
3 Jun 2026
Agents

AUGUSTE: Online-Learning dApp for Predictive URLLC Scheduling

DGX agent

arXiv:2606.03664v1 Announce Type: cross Abstract: Ultra Reliable and Low Latency Communications (URLLC) was one of the main motivations behind 5G, with 3GPP advertising 1-10 ms latency targets for app

agentsarxiv-cs-ai
3 Jun 2026
Research

CAD-to-CT Registration of Cylindrical Objects via Ellipse-Based Axis Estimation

DGX agent

arXiv:2606.02935v1 Announce Type: new Abstract: Accurate registration of CAD models to CT scans is essential for establishing ground truth geometry in volumetric imaging. Obtaining reliable object mas

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Calibration Data Trade-offs Across Capability Dimensions: Why Multi-Source Mixing Matters for High-Sparsity LLM Pruning

DGX agent

arXiv:2606.03328v1 Announce Type: cross Abstract: Post-training pruning compresses large language models to high sparsity using a small unlabelled calibration set, and recent work has concluded that t

model-releasesarxiv-cs-ai
3 Jun 2026
Local Ai

DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

DGX agent

arXiv:2606.03601v1 Announce Type: cross Abstract: While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rej

local-aiarxiv-cs-ai
3 Jun 2026
Model Releases

Decoupled Smart Contract Audits: Lightweight LLM Framework via Distillation and Aggregation

DGX agent

arXiv:2606.03128v1 Announce Type: cross Abstract: Smart contracts face critical security challenges that require thorough auditing in decentralized web services. While Large Language Models (LLMs) hav

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition

DGX agent

arXiv:2606.03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a functi

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

DGX agent

arXiv:2606.03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identifica

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

DGX agent

arXiv:2606.02971v1 Announce Type: new Abstract: Extracting reporting obligations from EU legislation is critical for assessing and reducing regulatory reporting burden. However, distinguishing reporti

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

DGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

model-releasesarxiv-cs-ai
3 Jun 2026
Agents

FORGE: Multi-Agent Graduated Exploitation and Detection Engineering

DGX agent

arXiv:2606.03453v1 Announce Type: cross Abstract: Vulnerability disclosure volumes now far exceed organizational assessment capacity, yet three adjacent research communities (proof-of-concept generati

agentsarxiv-cs-ai
3 Jun 2026
Model Releases

GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation

DGX agent

arXiv:2606.03682v1 Announce Type: new Abstract: Embodied navigation connects intelligent agents with the physical world and is fundamental for general robotic intelligence. Limited availability and qu

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

DGX agent

arXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as

model-releasesarxiv-cs-ai
3 Jun 2026
Research

High-Precision APT Malware Attribution with Out-of-Scope Resilience

DGX agent

arXiv:2606.03523v1 Announce Type: cross Abstract: Early attribution of Advanced Persistent Threat (APT) activity can help defenders prioritise investigation, select countermeasures, and reduce the imp

researcharxiv-cs-ai
3 Jun 2026
Model Releases

How Many Trees in a Random Forest? A Revisited Approach with Plateau Search and Optuna Integration

DGX agent

arXiv:2606.03549v1 Announce Type: new Abstract: Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves m

model-releasesarxiv-cs-lg
3 Jun 2026
Research

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems

DGX agent

arXiv:2606.02588v1 Announce Type: cross Abstract: We present Lean-GAP (Lean-Graduate Agebra Problems), 430 formalized graduate-level algebra problems from the textbook Abstract Algebra by Dummit and F

researcharxiv-cs-ai
3 Jun 2026
Model Releases

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

DGX agent

arXiv:2606.03303v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance

DGX agent

arXiv:2511.10055v2 Announce Type: replace Abstract: The performance of image generation has been significantly improved in recent years. However, the study of image screening is rare, and its performa

safetyarxiv-cs-cv
3 Jun 2026
Model Releases

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

DGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

model-releasesarxiv-cs-ai
3 Jun 2026
Tutorials

SAIL: Sound Abstract Interpreters with LLMs

DGX agent

arXiv:2511.13663v2 Announce Type: replace-cross Abstract: How to construct globally sound abstract interpreters to safely approximate program behaviors remains a bottleneck in abstract interpretation.

tutorialsarxiv-cs-lg
3 Jun 2026
Model Releases

SLU-2K: A Question-Based Benchmark for Semantic Evaluation of Sign Language Translation

DGX agent

arXiv:2606.03788v1 Announce Type: new Abstract: Sign Language Translation (SLT) is typically evaluated with surface-form metrics such as BLEU and ROUGE, which reward lexical overlap but do not directl

model-releasesarxiv-cs-cv
3 Jun 2026
Research

Social Caption: Evaluating Social Understanding in Multimodal Models

DGX agent

arXiv:2601.14569v2 Announce Type: replace Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions. We introduce SOCIAL

researcharxiv-cs-cl
3 Jun 2026
Agents

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

DGX agent

arXiv:2606.03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to bu

agentsarxiv-cs-ai
3 Jun 2026
← Previous
1…5859606162…82
Next →