AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,343
  • Agents7,552
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,164
  • Local Ai4,928
  • Model Releases23,861
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,402

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,343
  • Agents7,552
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,164
  • Local Ai4,928
  • Model Releases23,861
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,402

Source
Human
88,343Total entries
1Added by human
88,342Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
88,342 results
30 Jun 2026

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

SafetyDGX agent

arXiv:2606.29689v1 Announce Type: new Abstract: Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): unlike multiple-choice aesthetic benchmarks, it has no single

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

Model ReleasesDGX agent

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

Can regularization based JEPA (e.g. SIGReg) scale and compete with SOTA foundation models (DINO)? Here is the answer: yes and with 10x less …

Research
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

Can regularization based JEPA (e.g. SIGReg) scale and compete with SOTA foundation models (DINO)? Here is the answer: yes and with 10x less data. VISReg (slight variation of SIGReg) competes with DINO

CAN We Trust Your Results? A Cross-Dataset Study of Automotive IDS Evaluation

ResearchDGX agent

arXiv:2606.30430v1 Announce Type: cross Abstract: The increasing connectivity of modern vehicles has made securing in-vehicle communication networks a critical challenge. Intrusion Detection Systems (

Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks

AgentsDGX agent

arXiv:2606.28679v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly read untrusted content while holding side-effecting tools such as payments, email, CRM, and infrastructure APIs, ye

CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training

AgentsDGX agent

arXiv:2603.23559v2 Announce Type: replace-cross Abstract: GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screenshots

CAR: Cross-Vehicle Kinodynamics Adaptation via Mobility Representation

AgentsDGX agent

arXiv:2603.06866v3 Announce Type: replace Abstract: Developing autonomous mobile robot systems typically requires either extensive, platform-specific data collection or relies on simplified abstractio

CAREBench: A Child-Safety Risk Benchmark for Language Models

Model ReleasesDGX agent

arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations

CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models

SafetyDGX agent

arXiv:2606.30236v1 Announce Type: new Abstract: Medication errors, particularly dosing errors in clinical trials (CT), can lead to patient harm, adverse drug events and worse patient outcomes. Dosing

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance

SafetyDGX agent

arXiv:2606.28360v1 Announce Type: cross Abstract: University students often struggle to navigate complex academic policies, leading to advising bottlenecks and delayed access to critical information.

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety ben

Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition

ResearchDGX agent

arXiv:2601.06972v2 Announce Type: replace Abstract: In speech language modeling, two architectures dominate the frontier: the Transformer and the Conformer. However, it remains unknown whether their c

Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch

SafetyDGX agent

arXiv:2606.28815v1 Announce Type: cross Abstract: Mathswitch is an open-source project that imports mathematical concept records from sources such as Wikidata, Wikipedia, MathWorld, Encyclopedia of Ma

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

Model ReleasesDGX agent

arXiv:2406.08311v3 Announce Type: replace-cross Abstract: Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate

CaveAgent: Transforming LLMs into Stateful Runtime Operators

AgentsDGX agent

arXiv:2601.01569v4 Announce Type: replace Abstract: LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradigms that s

CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation

Local AiDGX agent

arXiv:2606.28724v1 Announce Type: cross Abstract: Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing. However, traditional I

CellDETR: A Detection-Guided Framework for Scalable Cell Representation Learning from Histopathology Images

ResearchDGX agent

arXiv:2606.29463v1 Announce Type: new Abstract: Recent advances in pathology foundation models have substantially improved patch and slide level representation learning from whole-slide images (WSIs).

Certified token billionaire. Having a blast at AIE, fantastic conf @swyx

ToolsDGX agent

The post references someone attending AIE (likely an AI or tech conference) and describes having a positive experience, with 'certified token billionaire' appearing to be either a humorous self-descri

Chamber geometry and specification numbers of Boolean threshold functions

ResearchDGX agent

arXiv:2606.29477v1 Announce Type: cross Abstract: The specification number sigma_n(f) of a Boolean threshold function f on n variables is the least number of points whose f-values determine f uniquely

Character Recognition of Nepali Number Plate

ApplicationsDGX agent

arXiv:2606.28946v1 Announce Type: new Abstract: This paper presents a robust Automatic Number Plate Recognition (ANPR) system tailored for Nepali license plates written in Devanagari script. In this p

Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem

SafetyDGX agent

arXiv:2606.29116v1 Announce Type: new Abstract: Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflows that combin

Characterizing Optimizer-Dependent Training Dynamics Through Hessian Eigenvector Displacement and Localization

Local AiDGX agent

arXiv:2606.30226v1 Announce Type: new Abstract: Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimiza

@charlieholtz preach! “Factories” is a depressing vision of the future, metaphors matter

ToolsDGX agent

This post discusses how the metaphor of 'factories' for AI systems presents a pessimistic framing of the future, arguing that the language and analogies we use to describe AI development shape our per

Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models

ResearchDGX agent

arXiv:2606.29897v1 Announce Type: cross Abstract: Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization systems are

China’s Meituan open-sources massive LongCat-2.0 AI model, saying it was trained on domestic chips

Model ReleasesDGX agent

Beijing, China-based Meituan Inc. today debuted its next-generation LongCat-2.0 open-source large language model, stating that the company trained the 1.6-trillion-parameter model on domestic Chinese

Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation

AgentsDGX agent

arXiv:2602.12089v3 Announce Type: replace-cross Abstract: As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove bot

Chronos: A Physics-Informed Full-History Framework for Non-Markovian Long-Horizon Manipulation

SafetyDGX agent

arXiv:2606.30318v1 Announce Type: new Abstract: General-purpose robot policies should be modeled as dynamical systems, yet many VLA and generative imitation policies still rely on present observations

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts

TutorialsDGX agent

arXiv:2510.09278v2 Announce Type: replace-cross Abstract: Training expert LLMs in domains with scarce data is difficult, often relying on multiple-choice questions (MCQs). However, standard outcome-ba

Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration

AgentsDGX agent

arXiv:2606.30246v1 Announce Type: new Abstract: Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant tas

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop …

Model ReleasesDGX agent

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop experience with Claude Code, Claude Cowork, and chat on all

Claude Science is Anthropic’s newest flagship product

Model ReleasesDGX agent

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way

Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively (Zac Hall/9to5Mac)

Model ReleasesDGX agent

Zac Hall / 9to5Mac: Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively — Anthropic is upgrading Claude Sonnet,

Claude Sonnet 5 is now available in Cursor. On CursorBench, it's a meaningful step up from Sonnet 4.6: 57% vs. 49%.

Model ReleasesDGX agent

Claude Sonnet 5 is now available as a model option in the Cursor code editor. On CursorBench, Sonnet 5 achieves a 57% score compared to Sonnet 4.6's 49%, representing a meaningful performance improvem

Claude Sonnet 5 is now available in Devin Desktop and Devin CLI. Sonnet 5 pairs frontier-level coding performance with a more affordable pri…

Model ReleasesDGX agent

Claude Sonnet 5 has been integrated into Devin Desktop and Devin CLI, offering advanced coding capabilities at a more competitive price point than previous models. This release represents an update to

Claude Sonnet 5 is now available in Perplexity for Pro and Max subscribers. You can also select it as an orchestrator model in Computer.

Model ReleasesDGX agent

Claude Sonnet 5 is now available for use in Perplexity, accessible to Pro and Max tier subscribers. Users can also select Claude Sonnet 5 as an orchestrator model within Perplexity's Computer feature,

Claude Sonnet 5 now available on Vercel AI Gateway

Model ReleasesDGX agent

Claude Sonnet 5 is now accessible through Vercel's AI Gateway, enabling developers to integrate Anthropic's latest language model into their applications via Vercel's unified API platform. This integr

CLEAR-MoE: Shared-Basis Expert Extraction from Frozen Vision Transformers via Calibration-Driven Layer Selection

HardwareDGX agent

arXiv:2606.28516v1 Announce Type: new Abstract: We present CLEAR-MoE, a four-phase post-training pipeline that converts a frozen pretrained Vision Transformer (ViT) into a sparse Mixture-of-Experts (M

Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation

SafetyDGX agent

arXiv:2606.29805v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, ca

CLIMP: Contrastive Language-Image Mamba Pretraining

ResearchDGX agent

arXiv:2601.06891v2 Announce Type: replace Abstract: Contrastive Language-Image Pre-training (CLIP) relies on Vision Transformers whose attention mechanism is susceptible to spurious correlations, and

Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

ResearchDGX agent

arXiv:2606.29876v1 Announce Type: cross Abstract: Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot distinguish stable

Clinical Risk-Aware Multi-Level Grading for Coronary Artery Stenosis through Curved Feature Reconstruction

ResearchDGX agent

arXiv:2606.30082v1 Announce Type: new Abstract: Developing a multi-level grading model for coronary artery stenosis holds great clinical significance for the diagnosis of coronary artery disease. Howe

CLMASP: Coupling Large Language Models with Answer Set Programming for Robotic Task Planning

ResearchDGX agent

arXiv:2406.03367v2 Announce Type: replace Abstract: Large Language Models (LLMs) possess extensive foundational knowledge and moderate reasoning abilities, making them suitable for general task planni

Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks

ResearchDGX agent

arXiv:2606.28662v1 Announce Type: cross Abstract: The flatness hypothesis suggests that flatness of the loss landscape, as measured by the eigenvalues of the loss Hessian, correlates with better neura

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2606.28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructi

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

Model ReleasesDGX agent

arXiv:2606.29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

Model ReleasesDGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System

TutorialsDGX agent

arXiv:2606.28953v1 Announce Type: cross Abstract: Poisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack scenario on a

Clustering with Non-adaptive Subset Queries

ResearchDGX agent

arXiv:2409.10908v3 Announce Type: replace-cross Abstract: Recovering the underlying k-clustering of a set U of n points by asking pair-wise same-cluster queries has garnered significant interest in th

ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation

Local AiDGX agent

arXiv:2512.02453v2 Announce Type: replace Abstract: Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and i

CMSL: Constructive Multi-Sequence Learning for Recommendation Systems

ResearchDGX agent

arXiv:2606.28533v1 Announce Type: cross Abstract: Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Models (DLRM) b

CMTFormer: Marrying Transformer with Hierarchical Information Interaction for RGB-Event Object Detection

SafetyDGX agent

arXiv:2606.29136v1 Announce Type: cross Abstract: Event cameras capture sparse brightness changes with high temporal resolution and high dynamic range, compensating for the deficiencies of the convent

Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action

AgentsDGX agent

arXiv:2506.13932v3 Announce Type: replace-cross Abstract: The rise of large language models (LLMs) has led to dramatic improvements across a wide range of natural language tasks. Their performance on

Cognitive World Models for Process-Level Social Influence Evaluation

Model ReleasesDGX agent

arXiv:2606.29495v1 Announce Type: new Abstract: Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, de

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

Model ReleasesDGX agent

arXiv:2606.28820v1 Announce Type: new Abstract: Reconstructing dynamic human--object interaction scenes from monocular video is difficult because the human, manipulated object, and background obey dif

CogSENet: Blind Image Deblurring with Blur-Conditioned Semantic Routing and Explicit Frequency Fusion

ApplicationsDGX agent

arXiv:2606.30030v1 Announce Type: new Abstract: Blind image deblurring demands the recovery of high-fidelity details and coherent structures from complex, unknown degradations. Current blind image deb

COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies

Model ReleasesDGX agent

arXiv:2606.30479v1 Announce Type: cross Abstract: Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adver

CollabOD: Collaborative Multi-Backbone with Cross-scale Vision for UAV Small Object Detection

Local AiDGX agent

arXiv:2603.05905v2 Announce Type: replace Abstract: Small object detection in unmanned aerial vehicle (UAV) imagery is challenging because high-altitude viewpoints produce severe scale variation, weak

Collective cooperation without individual fidelity in LLM agents

Model ReleasesDGX agent

arXiv:2606.30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be inter

CoLR-Det: Collaborative Latent Restoration for Small Object Detection in Low-Resolution Remote Sensing Images

ResearchDGX agent

arXiv:2601.12507v2 Announce Type: replace Abstract: Low-resolution remote sensing small object detection is limited by both missing visual details and the ambiguity of how details serve detection. Exi

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.28719v1 Announce Type: new Abstract: Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, exist

← Previous
1…439440441442443…1473
Next →