AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,770 results
Model Releases

AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

DGX agent

arXiv:2606.03116v1 Announce Type: cross Abstract: The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated eval

model-releasesarxiv-cs-ai
3 Jun 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

APIC: Amortized Physics-Informed Calibration using Neural Processes

DGX agent

arXiv:2606.03355v1 Announce Type: new Abstract: Physics models are inherently imperfect due to misspecified or missing mechanisms, resulting in systematic discrepancies between model predictions and r

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

ArrowFlow: Hierarchical Machine Learning in the Space of Permutations

DGX agent

arXiv:2604.04087v2 Announce Type: replace Abstract: We introduce ArrowFlow, a machine learning architecture that operates entirely in the space of permutations. Its computational units are ranking fil

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

As AI gets better, it reveals an empty promise

DGX agent

This week we've got tandem hands-ons with Google's new Gemini AI agent - Spark - from my colleagues David Pierce and Jay Peters. Their takeaways are similar: It's so effective that it's scary. Spark k

model-releasesthe-verge-ai
3 Jun 2026
Model Releases

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation

DGX agent

arXiv:2606.03175v1 Announce Type: new Abstract: Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an underspecified natural-language d

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Assessing and Mitigating Miscalibration in LLM-Based Social Science Measurement

DGX agent

arXiv:2605.11954v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables t

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics

DGX agent

arXiv:2507.21638v2 Announce Type: replace Abstract: The development of reinforcement learning (RL) algorithms has been largely driven by ambitious challenge tasks and benchmarks. Games have dominated

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception

DGX agent

arXiv:2606.02924v1 Announce Type: new Abstract: Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and pot

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Auditable Climate Risk Intelligence from Fragmented ESG Data: Deterministic Orchestration and Imbalance-Aware Learning for Scope 1-3 Validation

DGX agent

arXiv:2606.02604v1 Announce Type: cross Abstract: ESG and climate risk data remain fragmented across heterogeneous Scope 1, Scope 2, and Scope 3 reporting environments, while conventional validation p

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification

DGX agent

arXiv:2606.03031v1 Announce Type: new Abstract: Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Auditing Engagement Incentives in the Kidfluencer Ecosystem: A Multimodal Weak Supervision Approach

DGX agent

arXiv:2606.03173v1 Announce Type: cross Abstract: The rise of `kidfluencers' on YouTube has raised ethical concerns about child digital labor and exploitation. While emerging legislation attempts to r

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

AURA: Action-Gated Memory for Robot Policies at Constant VRAM

DGX agent

arXiv:2606.02775v1 Announce Type: new Abstract: The KV-cache is the right memory for datacenters but the wrong memory for robots. Datacenter inference batches many short requests and resets them, amor

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

DGX agent

arXiv:2606.02809v1 Announce Type: new Abstract: Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation con

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

AVTrack: Audio-Visual Tracking in Human-centric Complex Scenes

DGX agent

arXiv:2606.02724v1 Announce Type: cross Abstract: Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

DGX agent

arXiv:2606.02798v1 Announce Type: new Abstract: Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited. Existing benchmarks

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Benchmarking Visual State Tracking in Multimodal Video Understanding

DGX agent

arXiv:2606.03920v1 Announce Type: new Abstract: Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacit

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

BEV-ODOM2: Enhanced BEV-based Monocular Visual Odometry with PV-BEV Fusion and Dense Flow Supervision for Ground Robots

DGX agent

arXiv:2509.14636v2 Announce Type: replace Abstract: Scale-consistent ego-motion estimation is fundamental for autonomous ground robots. Bird's-Eye-View (BEV) representation naturally addresses the sca

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

Beyond Encoder Accumulation: Measuring Encoder Roles in Multi-Encoder VLMs

DGX agent

arXiv:2606.03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions

DGX agent

arXiv:2606.03318v1 Announce Type: new Abstract: Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

DGX agent

arXiv:2606.03829v1 Announce Type: new Abstract: Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and a

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Calibration Data Trade-offs Across Capability Dimensions: Why Multi-Source Mixing Matters for High-Sparsity LLM Pruning

DGX agent

arXiv:2606.03328v1 Announce Type: cross Abstract: Post-training pruning compresses large language models to high sparsity using a small unlabelled calibration set, and recent work has concluded that t

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Can Factual Opinions Be Edited (Manipulated) in Large Language Models?

DGX agent

arXiv:2606.03096v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous. Cu

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams

DGX agent

arXiv:2603.19250v2 Announce Type: replace Abstract: Evaluating language models in streaming environments is critical, yet underexplored. Existing benchmarks either focus on single complex events or pr

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

CANMOT: Class-Aware Noise Modeling for Multi-Object Tracking in Autonomous Driving

DGX agent

arXiv:2606.03590v1 Announce Type: new Abstract: Kalman filter (KF)-based multi-object tracking (MOT) remains a strong baseline for autonomous driving due to its strong performance, computational effic

model-releasesarxiv-cs-ro
3 Jun 2026
Model Releases

CAPER: Clause-Aligned Process Supervision for Text-to-SQL

DGX agent

arXiv:2606.03327v1 Announce Type: cross Abstract: Text-to-SQL systems are typically evaluated by query-level execution correctness, but this terminal signal provides little guidance about which interm

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

Causal Neural Probabilistic Circuits

DGX agent

arXiv:2603.01372v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Causal Preference Elicitation

DGX agent

arXiv:2602.01483v2 Announce Type: replace-cross Abstract: We propose causal preference elicitation, a Bayesian framework for expert-in-the-loop causal discovery that actively queries local edge relati

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Characterizing Detectability in 3DGS Poisoning: A Stage-wise Benchmark

DGX agent

arXiv:2606.03499v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has rapidly emerged as a leading representation for real-time novel view synthesis, but recent work shows it is vulnerable

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Chatbots Output Meaningful (but Problematic) Language

DGX agent

arXiv:2606.02973v1 Announce Type: new Abstract: Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, 'What is the capital of Spain?' and Claude answers,

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

DGX agent

arXiv:2606.02802v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model struct

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

DGX agent

arXiv:2606.03157v1 Announce Type: new Abstract: Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

COD10K-C: Benchmarking Robustness of Camouflaged Object Detection Under Natural Image Corruptions

DGX agent

arXiv:2606.02603v1 Announce Type: new Abstract: Camouflaged object detection has improved substantially, but most standard benchmarks evaluate models only on clean images. This is not realistic becaus

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

DGX agent

arXiv:2606.03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks ca

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Coherent Swap Regret and Channel-Proof Learning

DGX agent

arXiv:2606.02655v1 Announce Type: cross Abstract: External regret certifies stability only against replacing one's behavior by a fixed alternative. In a quantum game, this misses a natural physical mo

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Combining Statistical Features and Deep Encodings for Rehearsal-Based Class-Incremental Time Series Classification

DGX agent

arXiv:2606.03292v1 Announce Type: cross Abstract: Many systems used in real-world environments require adding new categories and incorporating new information without forgetting what was previously le

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

CoMPAS3D: A Dataset and Benchmark for Interactive Motion

DGX agent

arXiv:2507.19684v2 Announce Type: replace-cross Abstract: Socially interactive humanoid robots must engage with humans through their bodies, adapting in real time to a partner's movement, intent, and

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Compress then Merge: From Multiple LoRAs into One Low-Rank Adapter

DGX agent

arXiv:2606.03723v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) enables parameter-efficient specialization of foundation models, but the proliferation of task-specific adapters fragments ca

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Consistent Yet Wrong: Evidence Insensitivity in Spatial Vision-Language Models

DGX agent

arXiv:2606.02742v1 Announce Type: new Abstract: Spatial reasoning is fundamental to robotics, autonomy, and embodied AI, yet modern vision-language models (VLMs) remain unreliable on metric distance q

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

CoralBay: A Self-Supervised CT Foundation Model

DGX agent

arXiv:2606.03888v1 Announce Type: new Abstract: Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-purpose visual representations that transfer effec

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Cosmos 3: Omnimodal World Models for Physical AI

DGX agent

arXiv:2606.02800v1 Announce Type: cross Abstract: We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Cost-Aware Query Routing in RAG: Empirical Analysis of Retrieval Depth Tradeoffs

DGX agent

arXiv:2606.02581v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) faces a fundamental three-way tension: deeper retrieval improves factual grounding but inflates token costs and e

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA

DGX agent

arXiv:2512.00360v2 Announce Type: replace Abstract: We study timestamped question answering over educational lecture videos under a single-GPU latency/memory budget. Given a natural-language query, th

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

CREward: A Type-Specific Creativity Reward Model

DGX agent

arXiv:2511.19995v2 Announce Type: replace Abstract: Creativity is a complex phenomenon. When it comes to representing and assessing creativity, treating it as a single undifferentiated quantity would

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Critical evaluation of PINN for FWD inverse analysis and differentiable FEM as an alternative

DGX agent

arXiv:2606.03210v1 Announce Type: cross Abstract: Automatic-differentiation-based inverse analysis methods, including physics-informed neural networks (PINNs) and differentiable programming, have rece

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing

DGX agent

arXiv:2606.03618v1 Announce Type: new Abstract: AI-assisted coding agents are bottlenecked by input-token cost. Two pathologies of raw human input drive much of this overhead: tokenization inefficienc

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

DGX agent

arXiv:2603.01576v3 Announce Type: replace Abstract: Geo-Foundation Models (GFMs) have been evaluated across diverse Earth observation task including multiple domains and have demonstrated strong poten

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data

DGX agent

arXiv:2606.03209v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often d

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Decoupled Smart Contract Audits: Lightweight LLM Framework via Distillation and Aggregation

DGX agent

arXiv:2606.03128v1 Announce Type: cross Abstract: Smart contracts face critical security challenges that require thorough auditing in decentralized web services. While Large Language Models (LLMs) hav

model-releasesarxiv-cs-ai
3 Jun 2026
← Previous
1…219220221222223…475
Next →