AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,577 results
30 Jun 2026

SVC-Probe: A Framework for Evaluating Perturbation Generalization in Spatial Foundation-Model Embeddings

Model ReleasesDGX agent

arXiv:2606.28465v1 Announce Type: cross Abstract: This work examines perturbation generalization in spatial foundation-model embeddings derived from fluorescence microscopy images. Although these mode

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

Model ReleasesDGX agent

arXiv:2603.12703v3 Announce Type: replace Abstract: Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video u

SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2511.06090v3 Announce Type: replace-cross Abstract: Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduce r

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Model ReleasesDGX agent

arXiv:2606.29957v1 Announce Type: cross Abstract: Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assi

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

Model ReleasesDGX agent

arXiv:2511.17649v4 Announce Type: replace-cross Abstract: Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyday h

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

Model ReleasesDGX agent

arXiv:2606.29171v1 Announce Type: cross Abstract: While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training dat

TextClusterLab: An Integrated Framework for Reliable Text Clustering Studies

Model ReleasesDGX agent

arXiv:2606.28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

Model ReleasesDGX agent

arXiv:2606.29575v1 Announce Type: cross Abstract: Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing o…

Model ReleasesDGX agent

Thank you to everyone to came to the Claude managed agents workshop at @aiDotEngineer with @gcemaj and I. We had an absolute blast sharing our journey and walking you through building your first agent

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

Model ReleasesDGX agent

arXiv:2606.29278v1 Announce Type: new Abstract: We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential

The Contagion Tensor: A Framework for Measuring Output-Distribution Coupling in Multi-Agent LLM Systems -- and Auditing the Claims It Enables

Model ReleasesDGX agent

arXiv:2606.28839v1 Announce Type: new Abstract: We introduce the Contagion Tensor, a measurement framework for quantifying how large language model (LLM) output distributions couple across modalities,

The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models

Model ReleasesDGX agent

arXiv:2606.29799v1 Announce Type: new Abstract: This project introduces the CRISTAL Method (Coherent Reliable Intentional Synthesis of Truthful Analysis Logic), a neurosymbolic framework for automatin

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

Model ReleasesDGX agent

arXiv:2606.28325v1 Announce Type: cross Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a se

The FIL Hypothesis: Inductive Biases Help with Kernel Engineering

Model ReleasesDGX agent

arXiv:2606.30442v1 Announce Type: new Abstract: The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in human knowle

The food delivery leader in China just dropped an open weights 1.6 trillion parameter model while half the USA was sleeping … 🫨 🇨🇳

Model ReleasesDGX agent

The food delivery leader in China just dropped an open weights 1.6 trillion parameter model while half the USA was sleeping … 🫨 🇨🇳 Meituan, China's largest food delivery platform, open-sourced a 1.6 t

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown th

The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding

Model ReleasesDGX agent

arXiv:2603.03305v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to generate executable outputs, JSON objects, and API calls, where a single syntax error ca

The Human Creativity Benchmark

Model ReleasesDGX agent

arXiv:2606.30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine di

The Interference Gap: Comparing Retrieval Bounds in Human Memory and RAG Systems

Model ReleasesDGX agent

arXiv:2606.28327v1 Announce Type: cross Abstract: How do retrieval bounds compare between human episodic memory and Retrieval-Augmented Generation (RAG) systems under semantic interference? We present

𝗚𝗟𝗠-𝟱.𝟮 (the latest open weights model) is having an Enterprise moment, and it is not an exaggeration.🚀 🔥 We have been impressed by h…

Model ReleasesDGX agent

𝗚𝗟𝗠-𝟱.𝟮 (the latest open weights model) is having an Enterprise moment, and it is not an exaggeration.🚀 🔥 We have been impressed by how strongly GLM-5.2 is pushing long-horizon performance .. not just

The NTNU System at the S&I Challenge 2025 SLA Open Track

Model ReleasesDGX agent

arXiv:2506.05121v3 Announce Type: replace Abstract: A recent line of research on spoken language assessment (SLA) employs neural models such as BERT and wav2vec 2.0 (W2V) to evaluate speaking proficie

The Verbose Context Problem in Medical Records

Model ReleasesDGX agent

arXiv:2606.29503v1 Announce Type: cross Abstract: The verbose context problem occurs when structured concepts have token-inefficient textual representations. This bottleneck is acute in population hea

this is what you look like with low rise pants

Model ReleasesDGX agent

This post likely showcases visual examples or commentary on the aesthetic appearance and fit of low-rise pants, a fashion trend that was particularly popular in the early 2000s and has experienced per

Thrilled to announce the Wearable AI Workshop at ECCV 2026 🎉 that we're organizing with an awesome group of folks across Meta Reality Labs,…

Model ReleasesDGX agent

Thrilled to announce the Wearable AI Workshop at ECCV 2026 🎉 that we're organizing with an awesome group of folks across Meta Reality Labs, AMI Labs, HKUST, Georgia Tech, UCF, and U. of Edinburgh. If

Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding

Model ReleasesDGX agent

arXiv:2601.04693v2 Announce Type: replace Abstract: Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scar

Toward an Energy-Optimized Operation of Data Centers Located in Wind Farms Using Reinforcement Learning

Model ReleasesDGX agent

arXiv:2606.30316v1 Announce Type: new Abstract: This paper studies Reinforcement Learning as an online controller for curtailment-aware workload shifting in wind-turbine-integrated high-performance co

Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback

Model ReleasesDGX agent

arXiv:2606.29700v1 Announce Type: new Abstract: Planning often requires symbolic specifications that are both executable and verifiable. For large language models deployed in autonomous or decision-su

Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural

Towards Generalizable and Evidential Nuclear Magnetic Resonance-Based Molecular Structure Elucidation via Large Language Model Agent

Model ReleasesDGX agent

arXiv:2606.29776v1 Announce Type: cross Abstract: Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra for unknown m

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

Model ReleasesDGX agent

arXiv:2606.30598v1 Announce Type: new Abstract: Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing lea

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

Model ReleasesDGX agent

arXiv:2512.13660v3 Announce Type: replace-cross Abstract: Spatial tracing, as a fundamental embodied interaction ability for robots, is inherently challenging as it requires multi-step metric-grounded

TraceLab: Characterizing Coding Agent Workloads for LLM Serving

Model ReleasesDGX agent

arXiv:2606.30560v1 Announce Type: cross Abstract: Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this challenge r

Training Vision-Language-Action Models with Dense Embodied Chain-of-Thought Supervision

Model ReleasesDGX agent

arXiv:2606.30552v1 Announce Type: cross Abstract: Cross-embodiment transfer in vision-language-action (VLA) models remains challenging because low-level state and action spaces differ fundamentally ac

Translating Natural Language to Strategic Temporal Specifications via LLMs

Model ReleasesDGX agent

arXiv:2606.30441v1 Announce Type: cross Abstract: A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS). However, writing corr

Travel-Oriented Reasoning Large Language Model via Domain-Specific Knowledge Graphs

Model ReleasesDGX agent

arXiv:2606.29254v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate broad reasoning abilities but struggle with accuracy and reliability in specialized domains such as travel, whe

TriageRA-CCF: Source-Side Clinical Confidence and Coverage Signals for Adaptive Rank Budgeting in Medical LLMs

Model ReleasesDGX agent

arXiv:2606.29375v1 Announce Type: new Abstract: Medical large language models are commonly adapted with a fixed low-rank budget, even though medical questions differ substantially in confidence, clini

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

Model ReleasesDGX agent

arXiv:2606.28480v1 Announce Type: cross Abstract: As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader ra

Tumor-aware augmentation with task-guided attention analysis improves rectal cancer segmentation from magnetic resonance images

Model ReleasesDGX agent

arXiv:2605.05522v2 Announce Type: replace-cross Abstract: Although self-supervised pretraining is expected to learn broadly transferable representations, its effectiveness across imaging modalities su

Tutorial on using Gemini live to build a voice agent Uses deepagents as a tool: offload complex work to this subagent, use Gemini live for t…

Model ReleasesDGX agent

Tutorial on using Gemini live to build a voice agent Uses deepagents as a tool: offload complex work to this subagent, use Gemini live for the naturalness/latency Building voice agents can come with t

Two kinds of robustness are not the same: disentangling fault tolerance and low-SNR robustness in multi-domain event detection on real data

Model ReleasesDGX agent

arXiv:2606.29339v1 Announce Type: cross Abstract: Reliable event detection underpins induced-seismicity monitoring for Carbon dioxide Capture and Storage (CCS) and geothermal operations, distributed a

UniCA: Bi-directional Cross-Attention with Positive Similarity Loss for Robust Multi-Modal Retrieval

Model ReleasesDGX agent

arXiv:2606.28350v1 Announce Type: cross Abstract: Multi-modal retrieval has become increasingly critical for handling the growing volume of integrated visual-textual data in real-world applications, b

Unified Enhancement of the Generalization and Robustness of Language Models via Bi-Stage Optimization

Model ReleasesDGX agent

arXiv:2503.16550v2 Announce Type: replace Abstract: Neural network language models (LMs) are confronted with significant challenges in generalization and robustness. Currently, many studies focus on i

Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature

Model ReleasesDGX agent

arXiv:2606.29667v1 Announce Type: cross Abstract: The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to

UrbanCDNet: Appearance-Robust and Boundary-Aware Bitemporal Change Detection for Korean Urban Building Monitoring

Model ReleasesDGX agent

arXiv:2606.29781v1 Announce Type: new Abstract: Urban building change detection from bi-temporal aerial imagery is important for redevelopment monitoring, infrastructure management, and unauthorized-c

Variance Reduction on the Camera Axis: Multi-View Score Distillation for 3D

Model ReleasesDGX agent

arXiv:2606.29964v1 Announce Type: new Abstract: Score distillation turns a pretrained 2D diffusion model into a 3D generator, but the per-step gradient is estimated from a single randomly chosen view:

VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection

Model ReleasesDGX agent

arXiv:2603.21526v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) offer a promising path toward interpretable deepfake detection by generating textual explanations. However,

ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models

Model ReleasesDGX agent

arXiv:2606.28804v1 Announce Type: new Abstract: Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluat

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On

Model ReleasesDGX agent

arXiv:2603.11734v2 Announce Type: replace Abstract: As virtual try-on (VTON) continues to advance, a growing number of real-world scenarios have emerged, pushing beyond the ability of the existing spe

We are making a deliberate effort to use GLM 5.2 with OpenCode internally at Jarvislabs. I have spoken to 3 enterprise customers last week, …

Model ReleasesDGX agent

We are making a deliberate effort to use GLM 5.2 with OpenCode internally at Jarvislabs. I have spoken to 3 enterprise customers last week, who are exploring to host multiple open source models and mo

We build differently than other tech firms. No superintelligence. No competition to spend the most money. We're creating empowering, efficie…

Model ReleasesDGX agent

We build differently than other tech firms. No superintelligence. No competition to spend the most money. We're creating empowering, efficient AI to enhance human potential, not replace it. With high

We just made it a lot easier for AI agents to work with your documents. LlamaParse MCP now does more than parse or classify files: it can pu…

Model ReleasesDGX agent

We just made it a lot easier for AI agents to work with your documents. LlamaParse MCP now does more than parse or classify files: it can pull structured data out of contracts, invoices, and reports a

We're excited to be sponsoring and co-organizing IOL-AI 2026, a new open challenge with the International Linguistics Olympiad @IOLing_offic…

Model ReleasesDGX agent

We're excited to be sponsoring and co-organizing IOL-AI 2026, a new open challenge with the International Linguistics Olympiad @IOLing_official.🌎💬 🌐 Open-science 🗓️ One-month competition 🎯 Targeting A

We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological …

Model ReleasesDGX agent

We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the right analysis path, and make judgment call

We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then in…

Model ReleasesDGX agent

We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then instantly animate them with the other—all at a fraction of the

We’ve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. We'll begin restoring acces…

Model ReleasesDGX agent

We’ve received notice that the Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5. We'll begin restoring access tomorrow, and will share an update soon. We’re grateful to

What an honor to emcee the first day of @aiDotEngineer and introduce the Software Factories Track Thank you @swyx & team, and @KeycardLabs f…

Model ReleasesDGX agent

What an honor to emcee the first day of @aiDotEngineer and introduce the Software Factories Track Thank you @swyx & team, and @KeycardLabs for the support. “A year ago @GeoffreyHuntley released the Ra

What Drives the Inlier-Memorization Effect? A Theory of Outlier Detection via Early Training Dynamics

Model ReleasesDGX agent

arXiv:2606.29791v1 Announce Type: cross Abstract: Outlier detection (OD) aims to identify anomalous instances by learning the underlying structure of normal data (inliers), and is particularly challen

What's new in Claude Sonnet 5

Model ReleasesDGX agent

What's new in Claude Sonnet 5 Claude Sonnet 5 came out this morning. I always head straight for the 'what's new' developer docs because they tend to have more actionable information than the official

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

Model ReleasesDGX agent

arXiv:2606.28438v1 Announce Type: cross Abstract: Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality control. We st

When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2606.28376v1 Announce Type: cross Abstract: Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and existing memor

← Previous
1…122123124125126…377
Next →