AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,566 results
30 Jun 2026

ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering

Model ReleasesDGX agent

arXiv:2606.29706v1 Announce Type: cross Abstract: Telecom question answering (QA) is a challenging setting for retrieval-augmented generation (RAG): evidence is fragmented across standards, papers, en

Assessing the Business Process Modeling Competences of Large Language Models

Model ReleasesDGX agent

arXiv:2601.21787v2 Announce Type: replace-cross Abstract: The creation of Business Process Model and Notation (BPMN) models is a complex and time-consuming task requiring both domain knowledge and pro

Attractor States Emerge in Multi-Turn LLM Conversations

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.30571v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in open-ended multi-agent settings, but the long-run dynamics of model--model interaction remain po

Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning

Model ReleasesDGX agent

arXiv:2510.12957v4 Announce Type: replace-cross Abstract: We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce extbf{Attribution Graphs} (AGs), whic

Audio-Visual Continual Test-Time Adaptation without Forgetting

Model ReleasesDGX agent

arXiv:2602.18528v2 Announce Type: replace Abstract: Audio-visual continual test-time adaptation involves continually adapting a source audio-visual model at test-time, to unlabeled non-stationary doma

AURORA: Asymmetry and Update-Induced Rotation for Robust Hallucination Detection in Large Language Models

Model ReleasesDGX agent

arXiv:2606.29545v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. However, their tendency

AWS launches an internal organization for AI-focused forward-deployed engineers, backed by $1B in resources, following OpenAI and others in launching FDE teams (Russell Brandom/TechCrunch)

Model ReleasesDGX agent

Russell Brandom / TechCrunch: AWS launches an internal organization for AI-focused forward-deployed engineers, backed by $1B in resources, following OpenAI and others in launching FDE teams — As compa

AWS launches forward-deployed engineering team to speed enterprise agentic AI adoption

Model ReleasesDGX agent

Amazon Web Services Inc. said today it’s rolling out a new dedicated organization to bring agentic artificial intelligence systems, built on the same technology, to customers by embedding engineers in

BaRA: Bayesian Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.29184v1 Announce Type: new Abstract: While Low-rank adaptation (LoRA) enables highly efficient fine-tuning by constraining task-specific updates to fixed low-rank subspaces, this rigid desi

Bayesian Best-Arm Identification with Abstention: A Polynomial-to-Exponential Phase Transition

Model ReleasesDGX agent

arXiv:2606.29203v1 Announce Type: new Abstract: We study the Bayesian fixed-budget best-arm identification problem in which a learner can abstain from making a terminal recommendation. Subject to an a

BEACON: A Bayesian Optimization Inspired Strategy for Efficient Novelty Search

Model ReleasesDGX agent

arXiv:2406.03616v5 Announce Type: replace-cross Abstract: Novelty search (NS) aims to uncover diverse system behaviors through simulation or experiment without requiring a pre-specified scalar objecti

Benchmark AUC Is Not Deployable Reliability: A Cross-Dataset Audit of Off-the-Shelf Features for Surveillance Video Anomaly Detection

Model ReleasesDGX agent

arXiv:2606.29506v1 Announce Type: new Abstract: Automated 'suspicious behavior' flagging is a headline promise of AI surveillance, and the field reports high frame-level ROC-AUC on standard video anom

Benchmarking Geospatial Foundation Models for Agriculture Applications

Model ReleasesDGX agent

arXiv:2606.29664v1 Announce Type: new Abstract: Geospatial foundation models pretrained on satellite imagery promise broad generalization across remote sensing tasks and regions, but their geographic

BERTomelo: Your Portuguese Encoder Best Friend

Model ReleasesDGX agent

arXiv:2606.28999v1 Announce Type: cross Abstract: Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multilingual models

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Model ReleasesDGX agent

arXiv:2606.30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This

Beyond IID: How General Are Tabular Foundation Models, Really?

Model ReleasesDGX agent

arXiv:2606.30410v1 Announce Type: cross Abstract: Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Research communi

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

Model ReleasesDGX agent

arXiv:2508.09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solvin

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

Model ReleasesDGX agent

arXiv:2606.28963v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributi

Beyond the Reranker: Do RAG Retrieval Enhancements Help Once a Strong Reranker Is Present?

Model ReleasesDGX agent

arXiv:2606.28367v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is routinely extended with methods meant to improve retrieval: query expansion, hierarchical and cross-document s

Beyond Trajectory Matching: Reflow with Marginal Distribution Alignment

Model ReleasesDGX agent

arXiv:2606.29287v1 Announce Type: cross Abstract: Diffusion and continuous-flow generative models achieve high-quality generation, and their deterministic sampling can be formulated as solving learned

Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs

Model ReleasesDGX agent

arXiv:2606.29860v1 Announce Type: new Abstract: Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incompleteness, knowle

Blackknife: Hard-Label Query-Limited Black-Box Attacks on Heterogeneous Graph Neural Networks

Model ReleasesDGX agent

arXiv:2606.29240v1 Announce Type: new Abstract: Heterogeneous graph neural networks (HGNNs) have achieved strong performance in modeling complex graph-structured data with multiple node and relation t

BrepLLM: Enabling Large Language Models to Understand Boundary Representations

Model ReleasesDGX agent

arXiv:2512.16413v2 Announce Type: replace Abstract: Current token-sequence-based Large Language Models (LLMs) struggle to directly process 3D Boundary Representation (B-rep) models that contain comple

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

Model ReleasesDGX agent

arXiv:2606.30049v1 Announce Type: new Abstract: Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monit

Bridging the NISQ and Fault-Tolerant Regimes: Generative-ML-Assisted Quantum Selected CI for Molecular Simulations

Model ReleasesDGX agent

arXiv:2606.30551v1 Announce Type: cross Abstract: Calculation of binding energies for protein-ligand molecular systems requires accurate treatment of the electronic structure, a quantum chemistry prob

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Model ReleasesDGX agent

arXiv:2606.29445v1 Announce Type: cross Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarka

Bringing speed and strong cost performance to the market with Gemini Omni Flash and Nano Banana 2 Lite

Model ReleasesDGX agent

Great creative happens when your tools move at the speed of your ideas. To help you create rich, reliable experiences while reducing regeneration time and costs, we’re adding two new models to Gemini

Build agents even faster with Gemini Enterprise Agent Platform’s fully-managed, remote MCP server

Model ReleasesDGX agent

A couple of months ago, we announced that over 50 Google-managed MCP servers are available. Today, we’ll dive into how to use the Gemini Enterprise Agent Platform remote MCP server to securely connect

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Model ReleasesDGX agent

arXiv:2606.28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity proble

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich archite…

Model ReleasesDGX agent

Building voice agents can come with tradeoffs. 💬 Better convos w/ speech to speech models Vs 🥪 More reliable harnesses w/ sandwich architecture How to build a voice research agent with both: ✅ Gemini

Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models

Model ReleasesDGX agent

arXiv:2606.28406v1 Announce Type: cross Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design sch

Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation

Model ReleasesDGX agent

arXiv:2511.05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning remain

Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

Model ReleasesDGX agent

arXiv:2606.29920v1 Announce Type: new Abstract: Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring. However, the reliab

Can LLMs Hire Fairly? Racial Bias in Resume Screening

Model ReleasesDGX agent

arXiv:2606.28978v1 Announce Type: new Abstract: We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (202

Can LLMs Prove Robotic Path Planning Optimality? A Benchmark for Research-Level Algorithm Verification

Model ReleasesDGX agent

arXiv:2603.19464v2 Announce Type: replace Abstract: Robotic path planning problems are often NP-hard, and practical solutions typically rely on approximation algorithms with provable performance guara

Can OCR-VLMs Read Devanagari? A Stress-Test Benchmark and Post-Correction Study

Model ReleasesDGX agent

arXiv:2606.29213v1 Announce Type: new Abstract: OCR systems, ranging from classical engines to specialised OCR vision-language models (OCR-VLMs) and frontier multimodal LLMs, report strong results on

CAREBench: A Child-Safety Risk Benchmark for Language Models

Model ReleasesDGX agent

arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety ben

Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework

Model ReleasesDGX agent

arXiv:2406.08311v3 Announce Type: replace-cross Abstract: Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivariate

China’s Meituan open-sources massive LongCat-2.0 AI model, saying it was trained on domestic chips

Model ReleasesDGX agent

Beijing, China-based Meituan Inc. today debuted its next-generation LongCat-2.0 open-source large language model, stating that the company trained the 1.6-trillion-parameter model on domestic Chinese

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop …

Model ReleasesDGX agent

Claude Desktop is now available on Linux (Ubuntu and Debian) in beta. Alongside the browser and terminal, you now get a first-class desktop experience with Claude Code, Claude Cowork, and chat on all

Claude Science is Anthropic’s newest flagship product

Model ReleasesDGX agent

At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way

Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively (Zac Hall/9to5Mac)

Model ReleasesDGX agent

Zac Hall / 9to5Mac: Claude Sonnet 5 costs 2 per 1M input tokens and 10 per 1M output tokens through August 31, after which prices rise to 3 and 15, respectively — Anthropic is upgrading Claude Sonnet,

Claude Sonnet 5 is now available in Cursor. On CursorBench, it's a meaningful step up from Sonnet 4.6: 57% vs. 49%.

Model ReleasesDGX agent

Claude Sonnet 5 is now available as a model option in the Cursor code editor. On CursorBench, Sonnet 5 achieves a 57% score compared to Sonnet 4.6's 49%, representing a meaningful performance improvem

Claude Sonnet 5 is now available in Devin Desktop and Devin CLI. Sonnet 5 pairs frontier-level coding performance with a more affordable pri…

Model ReleasesDGX agent

Claude Sonnet 5 has been integrated into Devin Desktop and Devin CLI, offering advanced coding capabilities at a more competitive price point than previous models. This release represents an update to

Claude Sonnet 5 is now available in Perplexity for Pro and Max subscribers. You can also select it as an orchestrator model in Computer.

Model ReleasesDGX agent

Claude Sonnet 5 is now available for use in Perplexity, accessible to Pro and Max tier subscribers. Users can also select Claude Sonnet 5 as an orchestrator model within Perplexity's Computer feature,

Claude Sonnet 5 now available on Vercel AI Gateway

Model ReleasesDGX agent

Claude Sonnet 5 is now accessible through Vercel's AI Gateway, enabling developers to integrate Anthropic's latest language model into their applications via Vercel's unified API platform. This integr

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2606.28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructi

Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense

Model ReleasesDGX agent

arXiv:2606.29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

Model ReleasesDGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

Cognitive World Models for Process-Level Social Influence Evaluation

Model ReleasesDGX agent

arXiv:2606.29495v1 Announce Type: new Abstract: Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, de

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

Model ReleasesDGX agent

arXiv:2606.28820v1 Announce Type: new Abstract: Reconstructing dynamic human--object interaction scenes from monocular video is difficult because the human, manipulated object, and background obey dif

COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies

Model ReleasesDGX agent

arXiv:2606.30479v1 Announce Type: cross Abstract: Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adver

Collective cooperation without individual fidelity in LLM agents

Model ReleasesDGX agent

arXiv:2606.30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be inter

ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.28719v1 Announce Type: new Abstract: Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, exist

Compositional Dynamics in Learning and Mechanics

Model ReleasesDGX agent

arXiv:2606.28984v1 Announce Type: cross Abstract: We give a single compositional setting in which gradient-based learning and Hamiltonian-style mechanics appear as functorial semantics. The syntax is

Compressed Sensing for Capability Localization in Large Language Models

Model ReleasesDGX agent

arXiv:2603.03335v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit a wide range of capabilities, including mathematical reasoning, code generation, and linguistic behaviors. We s

Conversational analytics in BigQuery brings trusted agentic reasoning to everyone

Model ReleasesDGX agent

Businesses run on fast decisions, but the teams who hold the answers are often buried under a backlog of routine requests, leaving users waiting in line for insights they need now. Today, we are bring

Conversational Query Engine for Mixed-Modality Heterogeneous Enterprise Data Sources

Model ReleasesDGX agent

arXiv:2606.28370v1 Announce Type: cross Abstract: Enterprise business intelligence queries span structured warehouses and unstructured document repositories -- modalities with fundamentally different

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level cod…

Model ReleasesDGX agent

Cool new paper from NVIDIA. Looks like agentic coding is moving into hardware design. HORIZON treats hardware design as repository-level code evolution. A Markdown harness becomes a project pack with

← Previous
1…116117118119120…377
Next →