AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,292 results
5 Aug 2026

DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

Model ReleasesDGX agent

arXiv:2608.03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to

Disentangling MLP Neuron Weights in Vocabulary Space

Model ReleasesDGX agent

arXiv:2604.06005v2 Announce Type: replace Abstract: Interpreting the information encoded in language model weights remains a fundamental challenge in mechanistic interpretability. In this work, we int

Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.03297v1 Announce Type: new Abstract: A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant infor

Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA

Model ReleasesDGX agent

arXiv:2608.03177v1 Announce Type: new Abstract: How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misclassification

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all…

Model ReleasesDGX agent

Document OCR is not Getting Commoditized (by Frontier Models) The most common question I get is whether frontier models are going to eat all document processing solutions - just screenshot the page an

Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation

Model ReleasesDGX agent

arXiv:2608.03791v1 Announce Type: new Abstract: Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpo

Don't Walk the Line: Boundary Guidance for Filtered Generation

Model ReleasesDGX agent

arXiv:2510.11834v3 Announce Type: replace-cross Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tun

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Model ReleasesDGX agent

arXiv:2608.02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language pr

Double Descent in Gradient Boosting Decision Trees via Split-Candidate Scaling

Model ReleasesDGX agent

arXiv:2608.03111v1 Announce Type: new Abstract: Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width. For gradient boosting decision trees (GBDTs)

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

Model ReleasesDGX agent

arXiv:2608.03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attribu

DS@GT-ARC at eRisk 2026 Task 3: Sparse, Semantic, and LLM Reranking for ADHD Symptom Sentences

Model ReleasesDGX agent

arXiv:2608.03883v1 Announce Type: new Abstract: This paper describes our submissions to eRisk 2026 Task 3, ADHD Symptom Sentence Ranking. The task requires systems to rank candidate Reddit sentences a

Dynamically Allocating Evaluation Effort for Model Ranking

Model ReleasesDGX agent

arXiv:2608.03437v1 Announce Type: new Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing m

EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

Model ReleasesDGX agent

arXiv:2608.03179v1 Announce Type: new Abstract: Controllable local editing of 3D assets requires precise target localization and appropriate visual guidance. However, existing methods lack a simple ye

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

Model ReleasesDGX agent

arXiv:2608.03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only re

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Model ReleasesDGX agent

arXiv:2608.03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained fro

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

Model ReleasesDGX agent

arXiv:2608.03480v1 Announce Type: new Abstract: The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational

Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment

Model ReleasesDGX agent

arXiv:2407.21311v2 Announce Type: replace-cross Abstract: Unsupervised domain adaptation (UDA) aims to mitigate domain shift, where the distribution of labeled source data differs from that of unlabel

Enhancing Tabular Learners with Context-Aware Semantic Embeddings

Model ReleasesDGX agent

arXiv:2608.03565v1 Announce Type: new Abstract: While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discre

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

Model ReleasesDGX agent

arXiv:2608.03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations

Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform

Model ReleasesDGX agent

arXiv:2608.03311v1 Announce Type: cross Abstract: Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs.

Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks

Model ReleasesDGX agent

arXiv:2608.03794v1 Announce Type: cross Abstract: Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administra

Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

Model ReleasesDGX agent

arXiv:2608.02616v1 Announce Type: cross Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synth

Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment

Model ReleasesDGX agent

arXiv:2608.02786v1 Announce Type: new Abstract: AI systems can fail silently. The failure propagates through training loops, evaluation pipelines, and production monitoring stacks until downstream har

Exploiting Separability in Multi-Scale Grey-Box Bayesian Optimization

Model ReleasesDGX agent

arXiv:2608.03045v1 Announce Type: new Abstract: We consider grey-box optimization problems where the decision variables naturally partition into black-box variables (as arguments to an expensive black

Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors

Model ReleasesDGX agent

arXiv:2511.15968v2 Announce Type: replace-cross Abstract: External validation of breast ultrasound segmentation models remains limited because internal train--test splits do not capture domain shifts

Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks

Model ReleasesDGX agent

arXiv:2608.03222v1 Announce Type: cross Abstract: Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. F

FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection

Model ReleasesDGX agent

arXiv:2608.03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks rem

FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs

Model ReleasesDGX agent

arXiv:2608.03852v1 Announce Type: new Abstract: This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource c

Feels like Google could have been the dominating force in AI by open-sourcing the frontier with Gemini, Veo, and Nano Banana. Instead, they …

Model ReleasesDGX agent

Feels like Google could have been the dominating force in AI by open-sourcing the frontier with Gemini, Veo, and Nano Banana. Instead, they kept them behind APIs for a few billion dollars in revenue.

FinVerse: Financial Time-Series Benchmark

Model ReleasesDGX agent

arXiv:2608.03259v1 Announce Type: cross Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become incre

FLARE: Few-shot Learning-based Adaptive Reflective Engine

Model ReleasesDGX agent

arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-

Forecasting Revenue with its Customer-Base Drivers: When and Why Coordination Helps

Model ReleasesDGX agent

arXiv:2608.02911v1 Announce Type: new Abstract: Revenue forecasts guide acquisition budgets, demand planning, and customer-based valuations, yet an aggregate forecast does not show whether change refl

... four years later, I got the new Claude Fable 5 to actually build the game https://x.com/simonw/status/2085089518223602058

Model ReleasesDGX agent

... four years later, I got the new Claude Fable 5 to actually build the game https://x.com/simonw/status/2085089518223602058 Four years ago today I tweeted about having GPT-3 and DALL-E come up with

FreqAdapt: Frequency-Adaptive Processing for RAW Object Detection

Model ReleasesDGX agent

arXiv:2608.03385v1 Announce Type: new Abstract: Existing object detection methods predominantly utilize sRGB inputs, which are compressed from RAW sensor data using Image Signal Processors (ISP) origi

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

Model ReleasesDGX agent

arXiv:2508.00955v3 Announce Type: replace-cross Abstract: Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive

From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities

Model ReleasesDGX agent

arXiv:2608.03585v1 Announce Type: new Abstract: Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interp

Frozen High-Resolution Inference for Cross-City Object Detection: An AI City Challenge 2026 Study

Model ReleasesDGX agent

arXiv:2608.03136v1 Announce Type: new Abstract: Cross-city object detection requires a detector trained in one city to generalize to an unlabeled target city. In AI City Challenge 2026 Track 6, we ana

Fusion-Poly: A Polyhedral Framework Based on Spatial-Temporal Fusion for 3D Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2603.08199v2 Announce Type: replace Abstract: LiDAR-camera 3D multi-object tracking (MOT) combines rich visual semantics with accurate depth cues to improve trajectory consistency and tracking r

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Model ReleasesDGX agent

arXiv:2608.03764v1 Announce Type: new Abstract: Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-ev

GENESIS: Towards Explainable Causal Discovery

Model ReleasesDGX agent

arXiv:2608.03868v1 Announce Type: cross Abstract: Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve stru

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

Model ReleasesDGX agent

arXiv:2608.03826v1 Announce Type: new Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations,

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

Model ReleasesDGX agent

arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool

GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits

Model ReleasesDGX agent

arXiv:2608.02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fa

GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs

Model ReleasesDGX agent

arXiv:2608.03270v1 Announce Type: cross Abstract: GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resol

Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)

Model ReleasesDGX agent

Ivan Mehta / TechCrunch: Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release — Hark, a startup

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting th…

Model ReleasesDGX agent

Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting this. New research releases DataSpace, a benchmark where data

Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights fr…

Model ReleasesDGX agent

Hey everyone! I just released a sub-6B sparse activation AI model which was built with a brand new architecture : fusion. I fused weights from @liquidai's LFM2.5-2.6B & @Alibaba_Qwen's Qwen3.6-35B-A3B

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

Model ReleasesDGX agent

arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a ho

How Closely Do LLM Reviews Align with Human Peer Review?

Model ReleasesDGX agent

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

HUKUKBERT: Domain-Specific Language Model for Turkish Law

Model ReleasesDGX agent

arXiv:2604.04790v2 Announce Type: replace Abstract: Natural language processing (NLP) advances have powered a generation of LegalTech systems, but Turkish law remains under-served by domain-specific d

HyperFL: Query-Adaptive Representation Learning for Software Fault Localization

Model ReleasesDGX agent

arXiv:2608.02967v1 Announce Type: cross Abstract: Software fault localization identifies the code locations responsible for reported issues and is a fundamental step toward automated debugging and pro

HyVIC: A Metric-Driven Spatio-Spectral Hyperspectral Image Compression Architecture Based on Variational Autoencoders

Model ReleasesDGX agent

arXiv:2603.26468v2 Announce Type: replace Abstract: The rapid growth of hyperspectral data archives in remote sensing (RS) necessitates effective compression methods for storage and transmission. Rece

I remember a time when 'flash' meant 32B

Model ReleasesDGX agent

I mean, Deepseek V4 Flash is an absolutely fantastic model, even though I can't run it on my machine it's so fascinating to see how it performs. Knowing that potentially it could be run at home is rea

I updated my localy run benchmark with DeepSeek V4 Flash 0731

Model ReleasesDGX agent

It's the purple cluster on the top left (the good corner...) I'm running the MXFP4 version from Bartoswski with Dspark at 1K t/s prefill and 90 t/s gen (average). I tried different sampling params, yo

In-Context Collapse in Vision-Language Models and How to Mitigate it?

Model ReleasesDGX agent

arXiv:2608.02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely as

In-Context Pure Exploration in Continuous Decision Spaces

Model ReleasesDGX agent

arXiv:2602.17976v2 Announce Type: replace-cross Abstract: In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to ident

Incident Report: unsanctioned agent behaviour during cyber testing

Model ReleasesDGX agent

Incident Report: unsanctioned agent behaviour during cyber testing It happened again. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running

Inkling-Small 276B-A12B at ~2.9 tok/s on <10gb memory

Model ReleasesDGX agent

A follow up to the launch of Mference, it now supports and runs Inkling-Small 276B-A12B. Inkling-Small (Thinking Machines, Apache 2.0), from the pipenetwork/Inkling-Small-MLX-4bit conversion: 276B tot

Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation

Model ReleasesDGX agent

arXiv:2608.02639v1 Announce Type: cross Abstract: Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at th

Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning

Model ReleasesDGX agent

arXiv:2608.03138v1 Announce Type: cross Abstract: Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background, gap identif

← Previous
1…2930313233…372
Next →