AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,593 results
Model Releases

Bayesian Best-Arm Identification with Abstention: A Polynomial-to-Exponential Phase Transition

DGX agent

arXiv:2606.29203v1 Announce Type: new Abstract: We study the Bayesian fixed-budget best-arm identification problem in which a learner can abstain from making a terminal recommendation. Subject to an a

model-releasesarxiv-cs-lg
30 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

BEACON: A Bayesian Optimization Inspired Strategy for Efficient Novelty Search

DGX agent

arXiv:2406.03616v5 Announce Type: replace-cross Abstract: Novelty search (NS) aims to uncover diverse system behaviors through simulation or experiment without requiring a pre-specified scalar objecti

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Benchmark AUC Is Not Deployable Reliability: A Cross-Dataset Audit of Off-the-Shelf Features for Surveillance Video Anomaly Detection

DGX agent

arXiv:2606.29506v1 Announce Type: new Abstract: Automated 'suspicious behavior' flagging is a headline promise of AI surveillance, and the field reports high frame-level ROC-AUC on standard video anom

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Benchmarking Geospatial Foundation Models for Agriculture Applications

DGX agent

arXiv:2606.29664v1 Announce Type: new Abstract: Geospatial foundation models pretrained on satellite imagery promise broad generalization across remote sensing tasks and regions, but their geographic

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

BERTomelo: Your Portuguese Encoder Best Friend

DGX agent

arXiv:2606.28999v1 Announce Type: cross Abstract: Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

DGX agent

arXiv:2606.30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Beyond IID: How General Are Tabular Foundation Models, Really?

DGX agent

arXiv:2606.30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Research communi

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

DGX agent

arXiv:2508.09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solvin

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

DGX agent

arXiv:2606.28963v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributi

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Beyond the Reranker: Do RAG Retrieval Enhancements Help Once a Strong Reranker Is Present?

DGX agent

arXiv:2606.28367v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is routinely extended with methods meant to improve retrieval: query expansion, hierarchical and cross-document s

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Beyond Trajectory Matching: Reflow with Marginal Distribution Alignment

DGX agent

arXiv:2606.29287v1 Announce Type: cross Abstract: Diffusion and continuous-flow generative models achieve high-quality generation, and their deterministic sampling can be formulated as solving learned

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs

DGX agent

arXiv:2606.29860v1 Announce Type: new Abstract: Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incompleteness, knowle

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Blackknife: Hard-Label Query-Limited Black-Box Attacks on Heterogeneous Graph Neural Networks

DGX agent

arXiv:2606.29240v1 Announce Type: new Abstract: Heterogeneous graph neural networks (HGNNs) have achieved strong performance in modeling complex graph-structured data with multiple node and relation t

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

BrepLLM: Enabling Large Language Models to Understand Boundary Representations

DGX agent

arXiv:2512.16413v2 Announce Type: replace Abstract: Current token-sequence-based Large Language Models (LLMs) struggle to directly process 3D Boundary Representation (B-rep) models that contain comple

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

DGX agent

arXiv:2606.30049v1 Announce Type: new Abstract: Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monit

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Bridging the NISQ and Fault-Tolerant Regimes: Generative-ML-Assisted Quantum Selected CI for Molecular Simulations

DGX agent

arXiv:2606.30551v1 Announce Type: cross Abstract: Calculation of binding energies for protein-ligand molecular systems requires accurate treatment of the electronic structure, a quantum chemistry prob

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

DGX agent

arXiv:2606.29445v1 Announce Type: cross Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarka

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Bringing speed and strong cost performance to the market with Gemini Omni Flash and Nano Banana 2 Lite

DGX agent

Great creative happens when your tools move at the speed of your ideas. To help you create rich, reliable experiences while reducing regeneration time and costs, we’re adding two new models to Gemini

model-releasesgoogle-cloud-ai
30 Jun 2026
Model Releases

Build agents even faster with Gemini Enterprise Agent Platform’s fully-managed, remote MCP server

DGX agent

A couple of months ago, we announced that over 50 Google-managed MCP servers are available. Today, we’ll dive into how to use the Gemini Enterprise Agent Platform remote MCP server to securely connect

model-releasesgoogle-cloud-ai
30 Jun 2026
Model Releases

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

DGX agent

arXiv:2606.28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity proble

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich archite…

DGX agent

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich architecture How to build a voice research agent with both: ✅ Gemini

model-releasesharrison-chase--x
30 Jun 2026
Model Releases

Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

DGX agent

arXiv:2606.28406v1 Announce Type: cross Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

DGX agent

arXiv:2511.05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remain

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

DGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Can LLMs Hire Fairly? Racial Bias in Resume Screening

DGX agent

arXiv:2606.28978v1 Announce Type: new Abstract: We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (202

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification

DGX agent

arXiv:2603.19464v2 Announce Type: replace Abstract: Robotic path planning problems are often NP-hard, and practical solutions typically rely on approximation algorithms with provable performance guara

model-releasesarxiv-cs-ro
30 Jun 2026
Model Releases

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

DGX agent

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

CAREBench: A Child-Safety Risk Benchmark for Language Models

DGX agent

arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

DGX agent

arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety ben

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

DGX agent

arXiv:2406.08311v3 Announce Type: replace-cross Abstract: Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

China’s Meituan open-sources massive LongCat-2.0 AI model, saying it was trained on domestic chips

DGX agent

Beijing, China-based Meituan Inc. today debuted its next-generation LongCat-2.0 open-source large language model, stating that the company trained the 1.6-trillion-parameter model on domestic Chinese

model-releasessiliconangle
30 Jun 2026
Model Releases

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop …

DGX agent

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop experience with Claude Code, Claude Cowork, and chat on all

model-releasesthariq--x
30 Jun 2026
Model Releases

Claude Science is Anthropic’s newest flagship product

DGX agent

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way

model-releasesmit-tech-review
30 Jun 2026
Model Releases

Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively (Zac Hall/9to5Mac)

DGX agent

Zac Hall / 9to5Mac: Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively — Anthropic is upgrading Claude Sonnet,

model-releasestechmeme
30 Jun 2026
Model Releases

Claude Sonnet 5 is now available in Cursor. On CursorBench, it's a meaningful step up from Sonnet 4.6: 57% vs. 49%.

DGX agent

Claude Sonnet 5 is now available as a model option in the Cursor code editor. On CursorBench, Sonnet 5 achieves a 57% score compared to Sonnet 4.6's 49%, representing a meaningful performance improvem

model-releasescursor--x
30 Jun 2026
Model Releases

Claude Sonnet 5 is now available in Devin Desktop and Devin CLI. Sonnet 5 pairs frontier-level coding performance with a more affordable pri…

DGX agent

Claude Sonnet 5 has been integrated into Devin Desktop and Devin CLI, offering advanced coding capabilities at a more competitive price point than previous models. This release represents an update to

model-releasescognition-ai--x
30 Jun 2026
Model Releases

Claude Sonnet 5 is now available in Perplexity for Pro and Max subscribers. You can also select it as an orchestrator model in Computer.

DGX agent

Claude Sonnet 5 is now available for use in Perplexity, accessible to Pro and Max tier subscribers. Users can also select Claude Sonnet 5 as an orchestrator model within Perplexity's Computer feature,

model-releasesperplexity--x
30 Jun 2026
Model Releases

Claude Sonnet 5 now available on Vercel AI Gateway

DGX agent

Claude Sonnet 5 is now accessible through Vercel's AI Gateway, enabling developers to integrate Anthropic's latest language model into their applications via Vercel's unified API platform. This integr

model-releasesvercel-blog
30 Jun 2026
Model Releases

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

DGX agent

arXiv:2606.28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructi

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

DGX agent

arXiv:2606.29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

DGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Cognitive World Models for Process-Level Social Influence Evaluation

DGX agent

arXiv:2606.29495v1 Announce Type: new Abstract: Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, de

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

DGX agent

arXiv:2606.28820v1 Announce Type: new Abstract: Reconstructing dynamic human--object interaction scenes from monocular video is difficult because the human, manipulated object, and background obey dif

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies

DGX agent

arXiv:2606.30479v1 Announce Type: cross Abstract: Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adver

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Collective cooperation without individual fidelity in LLM agents

DGX agent

arXiv:2606.30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be inter

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

DGX agent

arXiv:2606.28719v1 Announce Type: new Abstract: Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, exist

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Compositional Dynamics in Learning and Mechanics

DGX agent

arXiv:2606.28984v1 Announce Type: cross Abstract: We give a single compositional setting in which gradient-based learning and Hamiltonian-style mechanics appear as functorial semantics. The syntax is

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Compressed Sensing for Capability Localization in Large Language Models

DGX agent

arXiv:2603.03335v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We s

model-releasesarxiv-cs-cl
30 Jun 2026
← Previous
1…147148149150151…471
Next →