AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,585 results
Model Releases

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

DGX agent

arXiv:2606.28963v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributi

model-releasesarxiv-cs-cl
30 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Beyond the Reranker: Do RAG Retrieval Enhancements Help Once a Strong Reranker Is Present?

DGX agent

arXiv:2606.28367v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is routinely extended with methods meant to improve retrieval: query expansion, hierarchical and cross-document s

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Beyond Trajectory Matching: Reflow with Marginal Distribution Alignment

DGX agent

arXiv:2606.29287v1 Announce Type: cross Abstract: Diffusion and continuous-flow generative models achieve high-quality generation, and their deterministic sampling can be formulated as solving learned

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs

DGX agent

arXiv:2606.29860v1 Announce Type: new Abstract: Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incompleteness, knowle

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Blackknife: Hard-Label Query-Limited Black-Box Attacks on Heterogeneous Graph Neural Networks

DGX agent

arXiv:2606.29240v1 Announce Type: new Abstract: Heterogeneous graph neural networks (HGNNs) have achieved strong performance in modeling complex graph-structured data with multiple node and relation t

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

BrepLLM: Enabling Large Language Models to Understand Boundary Representations

DGX agent

arXiv:2512.16413v2 Announce Type: replace Abstract: Current token-sequence-based Large Language Models (LLMs) struggle to directly process 3D Boundary Representation (B-rep) models that contain comple

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

DGX agent

arXiv:2606.30049v1 Announce Type: new Abstract: Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monit

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Bridging the NISQ and Fault-Tolerant Regimes: Generative-ML-Assisted Quantum Selected CI for Molecular Simulations

DGX agent

arXiv:2606.30551v1 Announce Type: cross Abstract: Calculation of binding energies for protein-ligand molecular systems requires accurate treatment of the electronic structure, a quantum chemistry prob

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

DGX agent

arXiv:2606.29445v1 Announce Type: cross Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarka

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Bringing speed and strong cost performance to the market with Gemini Omni Flash and Nano Banana 2 Lite

DGX agent

Great creative happens when your tools move at the speed of your ideas. To help you create rich, reliable experiences while reducing regeneration time and costs, we’re adding two new models to Gemini

model-releasesgoogle-cloud-ai
30 Jun 2026
Model Releases

Build agents even faster with Gemini Enterprise Agent Platform’s fully-managed, remote MCP server

DGX agent

A couple of months ago, we announced that over 50 Google-managed MCP servers are available. Today, we’ll dive into how to use the Gemini Enterprise Agent Platform remote MCP server to securely connect

model-releasesgoogle-cloud-ai
30 Jun 2026
Model Releases

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

DGX agent

arXiv:2606.28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity proble

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich archite…

DGX agent

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich architecture How to build a voice research agent with both: ✅ Gemini

model-releasesharrison-chase--x
30 Jun 2026
Model Releases

Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

DGX agent

arXiv:2606.28406v1 Announce Type: cross Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

DGX agent

arXiv:2511.05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remain

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

DGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Can LLMs Hire Fairly? Racial Bias in Resume Screening

DGX agent

arXiv:2606.28978v1 Announce Type: new Abstract: We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (202

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification

DGX agent

arXiv:2603.19464v2 Announce Type: replace Abstract: Robotic path planning problems are often NP-hard, and practical solutions typically rely on approximation algorithms with provable performance guara

model-releasesarxiv-cs-ro
30 Jun 2026
Model Releases

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

DGX agent

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

CAREBench: A Child-Safety Risk Benchmark for Language Models

DGX agent

arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

DGX agent

arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety ben

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

DGX agent

arXiv:2406.08311v3 Announce Type: replace-cross Abstract: Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

China’s Meituan open-sources massive LongCat-2.0 AI model, saying it was trained on domestic chips

DGX agent

Beijing, China-based Meituan Inc. today debuted its next-generation LongCat-2.0 open-source large language model, stating that the company trained the 1.6-trillion-parameter model on domestic Chinese

model-releasessiliconangle
30 Jun 2026
Model Releases

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop …

DGX agent

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop experience with Claude Code, Claude Cowork, and chat on all

model-releasesthariq--x
30 Jun 2026
Model Releases

Claude Science is Anthropic’s newest flagship product

DGX agent

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way

model-releasesmit-tech-review
30 Jun 2026
Model Releases

Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively (Zac Hall/9to5Mac)

DGX agent

Zac Hall / 9to5Mac: Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively — Anthropic is upgrading Claude Sonnet,

model-releasestechmeme
30 Jun 2026
Model Releases

Claude Sonnet 5 is now available in Cursor. On CursorBench, it's a meaningful step up from Sonnet 4.6: 57% vs. 49%.

DGX agent

Claude Sonnet 5 is now available as a model option in the Cursor code editor. On CursorBench, Sonnet 5 achieves a 57% score compared to Sonnet 4.6's 49%, representing a meaningful performance improvem

model-releasescursor--x
30 Jun 2026
Model Releases

Claude Sonnet 5 is now available in Devin Desktop and Devin CLI. Sonnet 5 pairs frontier-level coding performance with a more affordable pri…

DGX agent

Claude Sonnet 5 has been integrated into Devin Desktop and Devin CLI, offering advanced coding capabilities at a more competitive price point than previous models. This release represents an update to

model-releasescognition-ai--x
30 Jun 2026
Model Releases

Claude Sonnet 5 is now available in Perplexity for Pro and Max subscribers. You can also select it as an orchestrator model in Computer.

DGX agent

Claude Sonnet 5 is now available for use in Perplexity, accessible to Pro and Max tier subscribers. Users can also select Claude Sonnet 5 as an orchestrator model within Perplexity's Computer feature,

model-releasesperplexity--x
30 Jun 2026
Model Releases

Claude Sonnet 5 now available on Vercel AI Gateway

DGX agent

Claude Sonnet 5 is now accessible through Vercel's AI Gateway, enabling developers to integrate Anthropic's latest language model into their applications via Vercel's unified API platform. This integr

model-releasesvercel-blog
30 Jun 2026
Model Releases

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

DGX agent

arXiv:2606.28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructi

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

DGX agent

arXiv:2606.29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

DGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Cognitive World Models for Process-Level Social Influence Evaluation

DGX agent

arXiv:2606.29495v1 Announce Type: new Abstract: Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, de

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

DGX agent

arXiv:2606.28820v1 Announce Type: new Abstract: Reconstructing dynamic human--object interaction scenes from monocular video is difficult because the human, manipulated object, and background obey dif

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies

DGX agent

arXiv:2606.30479v1 Announce Type: cross Abstract: Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adver

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Collective cooperation without individual fidelity in LLM agents

DGX agent

arXiv:2606.30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be inter

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

DGX agent

arXiv:2606.28719v1 Announce Type: new Abstract: Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, exist

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Compositional Dynamics in Learning and Mechanics

DGX agent

arXiv:2606.28984v1 Announce Type: cross Abstract: We give a single compositional setting in which gradient-based learning and Hamiltonian-style mechanics appear as functorial semantics. The syntax is

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Compressed Sensing for Capability Localization in Large Language Models

DGX agent

arXiv:2603.03335v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We s

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Conversational analytics in BigQuery brings trusted agentic reasoning to everyone

DGX agent

Businesses run on fast decisions, but the teams who hold the answers are often buried under a backlog of routine requests, leaving users waiting in line for insights they need now. Today, we are bring

model-releasesgoogle-cloud-ai
30 Jun 2026
Model Releases

Conversational Query Engine for Mixed-Modality Heterogeneous Enterprise Data Sources

DGX agent

arXiv:2606.28370v1 Announce Type: cross Abstract: Enterprise business intelligence queries span structured warehouses and unstructured document repositories -- modalities with fundamentally different

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level cod…

DGX agent

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level code evolution. A Markdown harness becomes a project pack with

model-releasesdair-ai--x
30 Jun 2026
Model Releases

Core dump epidemiology: fixing an 18-year-old bug

DGX agent

This article describes OpenAI's investigation and resolution of a longstanding bug in their data infrastructure that had persisted for 18 years, likely related to how core dumps (system memory snapsho

model-releasesopenai
30 Jun 2026
Model Releases

CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

DGX agent

arXiv:2606.30175v1 Announce Type: new Abstract: The continuous evolution of large language models drives escalating demands on data scale and quality, and as different training stages impose increasin

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

DGX agent

arXiv:2512.10342v3 Announce Type: replace Abstract: Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual deci

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

DGX agent

arXiv:2511.02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Counterfactual Residual Data Augmentation for Regression

DGX agent

arXiv:2606.28460v1 Announce Type: cross Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations. Inspir

model-releasesarxiv-cs-ai
30 Jun 2026
← Previous
1…146147148149150…471
Next →