AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,292 results
6 Aug 2026

Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference

Model ReleasesDGX agent

arXiv:2608.04904v1 Announce Type: new Abstract: Multilingual large language models exhibit substantial performance differences across languages, while existing adaptation methods often require paramet

STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

Model ReleasesDGX agent

arXiv:2608.04567v1 Announce Type: new Abstract: Event knowledge concerns who does what to whom. Psycholinguists use event-plausibility judgments to examine how this knowledge supports human language p

Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up e…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up etched onto silicon As models satisfice etching makes sense,

Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays

Model ReleasesDGX agent

arXiv:2608.04043v1 Announce Type: new Abstract: Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation learning has concentrated on optical sens

Talk2Sensors: 3D Visual Grounding in Autonomous Driving via Sensor-Adaptive Physical Cue Matching

Model ReleasesDGX agent

arXiv:2608.04568v1 Announce Type: new Abstract: As a key capability for embodied intelligence, 3D visual grounding (3DVG) has been predominantly studied in indoor scenes with RGB-D or point-cloud inpu

Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding

Model ReleasesDGX agent

arXiv:2608.04127v1 Announce Type: new Abstract: Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and con

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Model ReleasesDGX agent

arXiv:2608.05138v1 Announce Type: cross Abstract: Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval

Test, then Route: How Language Models Execute In-Context Conditional Rules Across Models and Languages

Model ReleasesDGX agent

arXiv:2608.04183v1 Announce Type: new Abstract: When a language model follows an in-context conditional rule such as 'if P(x) then A else B,' does it assemble a runtime circuit with one module that te

Text2GraphQuery-Bench: A Text to Graph Query Benchmark

Model ReleasesDGX agent

arXiv:2602.11745v2 Announce Type: replace Abstract: Graph models are fundamental to data analysis in domains rich with complex relationships. Unlike SQL, which benefits from a rel- atively unified sta

The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

Model ReleasesDGX agent

arXiv:2608.04355v1 Announce Type: new Abstract: Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boun

The death of SLMs?

Model ReleasesDGX agent

I love to see these impressive models coming out that compete with the giants from companies like Z.ai, Moonshot, Alibaba, etc. A win for the open source/weight community is always welcome. While I am

The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

Model ReleasesDGX agent

arXiv:2608.04463v1 Announce Type: new Abstract: Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strate

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Model ReleasesDGX agent

arXiv:2608.04589v1 Announce Type: cross Abstract: EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize

The Loss Does Not See the Basis, but Adam Does

Model ReleasesDGX agent

arXiv:2608.05136v1 Announce Type: new Abstract: Gradient descent on a factored model W = UV^op is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization,

The new GPT-5.6 Sol powers all chats for paid users, including Instant, creating one consistent experience. In our high-stakes factuality ev…

Model ReleasesDGX agent

The new GPT-5.6 Sol powers all chats for paid users, including Instant, creating one consistent experience. In our high-stakes factuality evaluation covering finance, medicine and law, the new GPT‑5.6

The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals

Model ReleasesDGX agent

arXiv:2608.04611v1 Announce Type: cross Abstract: Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable

The Yokai Learning Environment: Tracking Beliefs Over Space and Time

Model ReleasesDGX agent

arXiv:2508.12480v3 Announce Type: replace Abstract: The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordination (ZS

They almost catched up on Frontier performance, so now catching up on prices

Model ReleasesDGX agent

Users also report that the free version was significantly downgraded after the release of the new models this is very important for us when considering local hosting. A lot of people decided not to bu

Thinking with Anchors: Grounded and Efficient Document Reasoning

Model ReleasesDGX agent

arXiv:2608.04424v1 Announce Type: new Abstract: Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reaso

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice…

Model ReleasesDGX agent

This paper by researchers from MIT and Stanford finds that most people would be financially better off if they followed the financial advice of LLMs (GPT-5.2 & Gemini 3 Flash) But some people get a bi

This webinar is happening in 30 minutes, and that means there's still time to register! Following the discussion will be an open Q&A with ou…

Model ReleasesDGX agent

This webinar is happening in 30 minutes, and that means there's still time to register! Following the discussion will be an open Q&A with our Head of AI Education @Prof_OZ, and the @arizeai team. See

TIDE: A Physically Diverse 3D Turbulence Benchmark Dataset for Advancing Scientific Machine Learning

Model ReleasesDGX agent

arXiv:2608.04222v1 Announce Type: cross Abstract: Turbulence is a central testbed for machine learning on physical dynamics because its governing laws are known exactly. However, most existing studies

Toward Federated Large Language Models in Medicine: A Parameter-Efficient Framework for Privacy-Preserving, Multi-Institutional Adaptation

Model ReleasesDGX agent

arXiv:2601.22124v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly adapted for medical applications, but most are trained using data from a single institution because pr

Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control

Model ReleasesDGX agent

arXiv:2608.04732v1 Announce Type: cross Abstract: Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning

Model ReleasesDGX agent

arXiv:2608.05139v1 Announce Type: new Abstract: Long-horizon reasoning in recent LLMs demands that the model switch between distinct skills inside a reasoning chain, such as first doing a math derivat

Towards a satellite image manipulation and deepfake localization benchmark dataset

Model ReleasesDGX agent

arXiv:2608.04840v1 Announce Type: cross Abstract: Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. Highly realisti

Towards Trustworthy Hypergraph Neural Networks under Label Noise

Model ReleasesDGX agent

arXiv:2608.04377v1 Announce Type: cross Abstract: Hypergraph neural networks (HGNNs) have demonstrated remarkable capabilities in processing complex higher-order relationships. However, their performa

Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision

Model ReleasesDGX agent

arXiv:2608.04879v1 Announce Type: cross Abstract: Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is indepe

Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

Model ReleasesDGX agent

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

Model ReleasesDGX agent

arXiv:2608.04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost

Tropical Algebraic Geometry for Neuronal Representations: An Arakelov-Green Measure Based Descriptor for Graph Learning

Model ReleasesDGX agent

arXiv:2608.04460v1 Announce Type: cross Abstract: The quantitative analysis of 3D neuronal morphologies requires capturing both graph topology and spatial geometry. Current message-passing Graph Neura

Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

Model ReleasesDGX agent

arXiv:2607.19322v2 Announce Type: replace Abstract: Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric req

Uber burned through its 2026 AI coding budget in four months. Microsoft canceled most of its Claude Code licenses six months after rolling t…

Model ReleasesDGX agent

Uber burned through its 2026 AI coding budget in four months. Microsoft canceled most of its Claude Code licenses six months after rolling them out. The mechanics are simple: per-token cost keeps fall

UG-UMRE: Uncertainty-Guided Modality Augmentation and Distributional Calibration for Unified Multimodal Relation Extraction

Model ReleasesDGX agent

arXiv:2608.04949v1 Announce Type: cross Abstract: Unified Multimodal Relation Extraction (UMRE) aims to identify intra-modal and cross-modal relations between textual entities and visual objects. Howe

UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models

Model ReleasesDGX agent

arXiv:2608.04701v1 Announce Type: new Abstract: The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generati

Unsloth's Gemma 4 mmproj silently broke vision & audio on newer llama.cpp builds — anyone else hit this?

Model ReleasesDGX agent

So I had been building ScreenMind, kinda like local ai desktop assistant that uses Gemma 4 for screen analysis, voice memo transcription, and meeting transcription — all through llama-server. Everythi

Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

Model ReleasesDGX agent

arXiv:2608.04902v1 Announce Type: new Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio

We made an MCP for your phone. Your laptop is just half of your life, and the other half is in your phone. Now your agent gets the mobile sc…

Model ReleasesDGX agent

We made an MCP for your phone. Your laptop is just half of your life, and the other half is in your phone. Now your agent gets the mobile screen too. No connectors, no complex setup, it has access to

We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, $0.18/task - ARC-…

Model ReleasesDGX agent

We re-tested GPT-5.6 Luna from @OpenAI on ARC-AGI (Verified) following its recent 80% price reduction: - ARC-AGI-2: 59.6%, 0.18/task - ARC-AGI-1: 90.7%, 0.07/task The new results match Luna's original

We’re making better intelligence easier to access in ChatGPT for everyone: - GPT-5.6 Sol now powers both Instant and deep reasoning for Plus…

Model ReleasesDGX agent

We’re making better intelligence easier to access in ChatGPT for everyone: - GPT-5.6 Sol now powers both Instant and deep reasoning for Plus & Pro users, delivering more factual, focused responses. -

What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend

Model ReleasesDGX agent

arXiv:2608.04714v1 Announce Type: cross Abstract: Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are co

When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs

Model ReleasesDGX agent

arXiv:2608.04893v1 Announce Type: cross Abstract: Multi-agent LLM systems relay key--value caches instead of text and credit their gains to exchanged ``latent thoughts''. That credit is a claim about

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

Model ReleasesDGX agent

arXiv:2608.04726v1 Announce Type: new Abstract: Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usual

Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

Model ReleasesDGX agent

arXiv:2608.04613v1 Announce Type: new Abstract: Anomaly detection is a safety-critical machine learning problem with applications ranging from fraud detection to network intrusion prevention and indus

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

Model ReleasesDGX agent

arXiv:2608.04964v1 Announce Type: new Abstract: Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods s

Your agentic summer: No-cost lessons from Google experts to build and scale agents

Model ReleasesDGX agent

I’ve talked to developers, IT leaders, and builders who all ask the same question: How do we actually get agents into production? The answer isn't theoretical — it's hands-on. Whether it’s designing a

Zero-shot reasoning for simulating scholarly peer-review

Model ReleasesDGX agent

arXiv:2510.02027v2 Announce Type: replace Abstract: Scholarly publishing requires scalable scrutiny supported by auditable evidence. This paper presents a two-component benchmark of xPeer, the peer-re

5 Aug 2026

3DGSI-Assessor: A Large-Scale Dataset and An LMM-based Method for 3D Gaussian Splatting Image Quality Assessment

Model ReleasesDGX agent

arXiv:2608.03279v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become a dominant representation for real-time novel view synthesis (NVS), yet its storage footprint makes compression

A ultra-lightweight mini agent - zero framework and local/ollama first

Model ReleasesDGX agent

https://github.com/mohsinkaleem/agent-mini.git A minimal 3k lines, local-first AI agent you can actually understand and extend. Optimized for smaller local models like qwen 3.6 4b or 9b pip install ag

A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation

Model ReleasesDGX agent

arXiv:2608.02805v1 Announce Type: cross Abstract: In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, using short-for

Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models

Model ReleasesDGX agent

arXiv:2608.03112v1 Announce Type: cross Abstract: Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands of tokens pe

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

Model ReleasesDGX agent

arXiv:2608.03569v1 Announce Type: new Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in eval

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

Model ReleasesDGX agent

arXiv:2608.03166v1 Announce Type: new Abstract: Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and educatio

AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits

Model ReleasesDGX agent

arXiv:2608.03738v1 Announce Type: new Abstract: As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level defects with

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2608.03744v1 Announce Type: new Abstract: Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be

AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?

Model ReleasesDGX agent

arXiv:2608.03076v1 Announce Type: new Abstract: Multi-agent studies commonly place AI agents in predefined games, markets, or roles, making it difficult to distinguish endogenous economic organization

AI Security Leaderboard: Methodology, Results and Minimal Standard

Model ReleasesDGX agent

arXiv:2608.03070v1 Announce Type: cross Abstract: Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much pro

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

Model ReleasesDGX agent

arXiv:2608.03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

Model ReleasesDGX agent

arXiv:2608.03154v1 Announce Type: new Abstract: Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and

ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts

Model ReleasesDGX agent

arXiv:2608.03898v1 Announce Type: new Abstract: The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a signific

← Previous
1…2728293031…372
Next →