AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,106 results
13 Apr 2026

PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos

Model ReleasesDGX agent

arXiv:2604.08991v1 Announce Type: cross Abstract: Small object-centric spatial understanding in indoor videos remains a significant challenge for multimodal large language models (MLLMs), despite its

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

Model ReleasesDGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

Precomputing Multi-Agent Path Replanning using Temporal Flexibility

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2601.04884v2 Announce Type: replace Abstract: Executing a multi-agent plan can be challenging when an agent is delayed, because this typically creates conflicts with other agents. So, we need to

Prefer watching on @YouTube while scrolling through the comment section? We get it: http://youtu.be/b1Pvt072wKQ?si=NGPgm30ur1WFtKQS

Model ReleasesDGX agent

Google AI shared a post on X (formerly Twitter) directing followers to watch their content on YouTube, suggesting their video is also available on that platform for viewers who prefer that experience.

Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search

Model ReleasesDGX agent

arXiv:2604.08598v1 Announce Type: cross Abstract: Text-based person search faces inherent limitations due to data scarcity, driven by stringent privacy constraints and the high cost of manual annotati

Provable Post-Training Quantization: Theoretical Analysis of OPTQ and Qronos

Model ReleasesDGX agent

arXiv:2508.04853v2 Announce Type: replace-cross Abstract: Post-training quantization (PTQ) has become a crucial tool for reducing the memory and compute costs of modern deep neural networks, including

QARIMA: A Quantum Approach To Classical Time Series Analysis

Model ReleasesDGX agent

arXiv:2604.08277v2 Announce Type: replace-cross Abstract: We present a quantum-inspired ARIMA methodology that integrates quantum-assisted lag discovery with fixed-configuration variational quantum ci

QoS-QoE Translation with Large Language Model

Model ReleasesDGX agent

arXiv:2604.08703v1 Announce Type: cross Abstract: QoS-QoE translation is a fundamental problem in multimedia systems because it characterizes how measurable system and network conditions affect user-p

QuanBench+: A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation

Model ReleasesDGX agent

arXiv:2604.08570v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code generation, yet quantum code generation is still evaluated mostly within single frameworks

Quantisation Reshapes the Metacognitive Geometry of Language Models

Model ReleasesDGX agent

arXiv:2604.08976v1 Announce Type: new Abstract: We report that model quantisation restructures domain-level metacognitive efficiency in LLMs rather than degrading it uniformly. Evaluating Llama-3-8B-I

R2G: A Multi-View Circuit Graph Benchmark Suite from RTL to GDSII

Model ReleasesDGX agent

arXiv:2604.08810v1 Announce Type: new Abstract: Graph neural networks (GNNs) are increasingly applied to physical design tasks such as congestion prediction and wirelength estimation, yet progress is

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models

Model ReleasesDGX agent

arXiv:2511.19704v2 Announce Type: replace Abstract: Open-vocabulary semantic segmentation (OVSS) underpins many vision and robotics tasks that require generalizable semantic understanding. Existing ap

RansomTrack: A Hybrid Behavioral Analysis Framework for Ransomware Detection

Model ReleasesDGX agent

arXiv:2604.08739v1 Announce Type: cross Abstract: Ransomware poses a serious and fast-acting threat to critical systems, often encrypting files within seconds of execution. Research indicates that ran

Realism character reference is on the way. Please fill out the survey below to sign up and get notified when it launches. https://links.comf…

Model ReleasesDGX agent

ComfyUI is developing a realism character reference feature and is collecting user interest through a survey prior to its official launch. Users can sign up via the provided survey link to receive not

Reasoning in a Combinatorial and Constrained World: Benchmarking LLMs on Natural-Language Combinatorial Optimization

Model ReleasesDGX agent

arXiv:2602.02188v2 Announce Type: replace Abstract: While large language models (LLMs) have shown strong performance in math and logic reasoning, their ability to handle combinatorial optimization (CO

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

Model ReleasesDGX agent

arXiv:2602.11354v2 Announce Type: replace Abstract: The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus primarily on

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation

Model ReleasesDGX agent

arXiv:2510.17640v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable performance on complex tasks through imitation learning in recent robotic man

Retrieval Augmented Classification for Confidential Documents

Model ReleasesDGX agent

arXiv:2604.08628v1 Announce Type: cross Abstract: Unauthorized disclosure of confidential documents demands robust, low-leakage classification. In real work environments, there is a lot of inflow and

Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios

Model ReleasesDGX agent

arXiv:2509.20006v3 Announce Type: replace Abstract: With the large models easing the labor-intensive manipulation process, image manipulations in today's real scenarios often entail a complex manipula

Robust Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2604.08571v1 Announce Type: cross Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their underlying reasoning processes remain highly ov

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.09452v1 Announce Type: cross Abstract: Safety guarantees are a prerequisite to the deployment of reinforcement learning (RL) agents in safety-critical tasks. Often, deployment environments

SAGE: A Service Agent Graph-guided Evaluation Benchmark

Model ReleasesDGX agent

arXiv:2604.09285v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Ex

Scaling model size is hitting diminishing returns. The real gains are in orchestration. Our Co-Founder & Co-CEO @yshoham makes the case in a…

Model ReleasesDGX agent

Scaling model size is hitting diminishing returns. The real gains are in orchestration. Our Co-Founder & Co-CEO @yshoham makes the case in a rare long-form profile by @Calcalistech today. The man tryi

SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

Model ReleasesDGX agent

arXiv:2604.08988v1 Announce Type: new Abstract: Current LLM-based agents demonstrate strong performance in episodic task execution but remain constrained by static toolsets and episodic amnesia, faili

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2512.02231v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely a

See you this Wednesday at the Ollama Gemma Meetup! 💎

Model ReleasesDGX agent

Ollama announced a meetup focused on Gemma, Google's open-weight language model family, scheduled for Wednesday. The event likely brought together developers and AI enthusiasts to explore running Gemm

Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise

Model ReleasesDGX agent

arXiv:2604.09532v1 Announce Type: cross Abstract: Prompt learning is a parameter-efficient approach for vision-language models, yet its robustness under label noise is less investigated. Visual conten

Semantic Rate-Distortion for Bounded Multi-Agent Communication: Capacity-Derived Semantic Spaces and the Communication Cost of Alignment

Model ReleasesDGX agent

arXiv:2604.09521v1 Announce Type: cross Abstract: When two agents of different computational capacities interact with the same environment, they need not compress a common semantic alphabet differentl

SenBen: Sensitive Scene Graphs for Explainable Content Moderation

Model ReleasesDGX agent

arXiv:2604.08819v1 Announce Type: cross Abstract: Content moderation systems classify images as safe or unsafe but lack spatial grounding and interpretability: they cannot explain what sensitive behav

Sentiment Classification of Gaza War Headlines: A Comparative Analysis of Large Language Models and Arabic Fine-Tuned BERT Models

Model ReleasesDGX agent

arXiv:2604.08566v1 Announce Type: new Abstract: This study examines how different artificial intelligence architectures interpret sentiment in conflict-related media discourse, using the 2023 Gaza War

SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding

Model ReleasesDGX agent

arXiv:2507.20185v2 Announce Type: replace Abstract: Session history is a common way of recording user interacting behaviors throughout a browsing activity with multiple products. For example, if an us

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos

Model ReleasesDGX agent

arXiv:2604.09037v1 Announce Type: cross Abstract: Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but over

SimScale: Learning to Drive via Real-World Simulation at Scale

Model ReleasesDGX agent

arXiv:2511.23369v3 Announce Type: replace Abstract: Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-di

Skill-Conditioned Visual Geolocation for Vision-Language

Model ReleasesDGX agent

arXiv:2604.09025v1 Announce Type: cross Abstract: Vision-language models (VLMs) have shown a promising ability in image geolocation, but they still lack structured geographic reasoning and the capacit

Skip-Connected Policy Optimization for Implicit Advantage

Model ReleasesDGX agent

arXiv:2604.08690v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has proven effective in RLVR by using outcome-based rewards. While fine-grained dense rewards can theoretica

So the concern over Mythos and cybersecurity seems warranted.

Model ReleasesDGX agent

So the concern over Mythos and cybersecurity seems warranted. We conducted cyber evaluations of Claude Mythos Preview and found that it is the first model to complete an AISI cyber range end-to-end. 🧵

Sources: SoftBank, Sony, Honda, and six other Japanese companies launch a new AI company to develop a 1T-parameter foundation model for 'physical AI' by 2030 (Natsuki Yamamoto/Nikkei Asia)

Model ReleasesDGX agent

Natsuki Yamamoto / Nikkei Asia: Sources: SoftBank, Sony, Honda, and six other Japanese companies launch a new AI company to develop a 1T-parameter foundation model for “physical AI” by 2030 — TOKYO —

SPASM: Stable Persona-driven Agent Simulation for Multi-turn Dialogue Generation

Model ReleasesDGX agent

arXiv:2604.09212v1 Announce Type: new Abstract: Large language models are increasingly deployed in multi-turn settings such as tutoring, support, and counseling, where reliability depends on preservin

Spectral Geometry of LoRA Adapters Encodes Training Objective and Predicts Harmful Compliance

Model ReleasesDGX agent

arXiv:2604.08844v1 Announce Type: new Abstract: We study whether low-rank spectral summaries of LoRA weight deltas can identify which fine-tuning objective was applied to a language model, and whether

Spectral-Transport Stability and Benign Overfitting in Interpolating Learning

Model ReleasesDGX agent

arXiv:2604.08625v1 Announce Type: cross Abstract: We develop a theoretical framework for generalization in the interpolating regime of statistical learning. The central question is why highly overpara

SPP-SBL: Space-Power Prior Sparse Bayesian Learning for Block Sparse Recovery

Model ReleasesDGX agent

arXiv:2505.08518v2 Announce Type: replace-cross Abstract: The recovery of block-sparse signals with unknown structural patterns remains a fundamental challenge in structured sparse signal reconstructi

STIndex: A Context-Aware Multi-Dimensional Spatiotemporal Information Extraction System

Model ReleasesDGX agent

arXiv:2604.08597v1 Announce Type: cross Abstract: Extracting structured knowledge from unstructured data still faces practical limitations: entity and event extraction pipelines remain brittle, knowle

StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding

Model ReleasesDGX agent

arXiv:2604.09000v1 Announce Type: new Abstract: Vision agent memory has shown remarkable effectiveness in streaming video understanding. However, storing such memory for videos incurs substantial memo

Structured Uncertainty guided Clarification for LLM Agents

Model ReleasesDGX agent

arXiv:2511.08798v2 Announce Type: replace-cross Abstract: LLM agents with tool-calling capabilities often fail when user instructions are ambiguous or incomplete, leading to incorrect invocations and

🚨 SUPER GEMMA 4 26B UNCENSORED IS INSANE LLM WIZARD COOKING AGAIN @songjunkr Dropped SuperGemma4-26B-Uncensored GGUF v2 and it’s trending o…

Model ReleasesDGX agent

🚨 SUPER GEMMA 4 26B UNCENSORED IS INSANE LLM WIZARD COOKING AGAIN @songjunkr Dropped SuperGemma4-26B-Uncensored GGUF v2 and it’s trending on @huggingface🤗 This thing SMOKES the regular Gemma-4 26B: 🤯0

Task-agnostic Low-rank Residual Adaptation for Efficient Federated Continual Fine-Tuning

Model ReleasesDGX agent

arXiv:2505.12318v2 Announce Type: replace Abstract: Federated Parameter-Efficient Fine-Tuning (Fed-PEFT) enables lightweight adaptation of large pre-trained models in federated learning settings by up

TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice

Model ReleasesDGX agent

arXiv:2604.08948v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel in various general domains, they exhibit notable gaps in the highly specialized, knowledge-intensive, and legal

Temperature-Dependent Performance of Prompting Strategies in Extended Reasoning Large Language Models

Model ReleasesDGX agent

arXiv:2604.08563v1 Announce Type: cross Abstract: Extended reasoning models represent a transformative shift in Large Language Model (LLM) capabilities by enabling explicit test-time computation for c

Temporal Dropout Risk in Learning Analytics: A Harmonized Survival Benchmark Across Dynamic and Early-Window Representations

Model ReleasesDGX agent

arXiv:2604.08870v1 Announce Type: cross Abstract: Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heterogeneous proto

Text-Conditioned Multi-Expert Regression Framework for Fully Automated Multi-Abutment Design

Model ReleasesDGX agent

arXiv:2604.09047v1 Announce Type: new Abstract: Dental implant abutments serve as the geometric and biomechanical interface between the implant fixture and the prosthetic crown, yet their design relie

The AI Codebase Maturity Model: From Assisted Coding to Self-Sustaining Systems

Model ReleasesDGX agent

arXiv:2604.09388v1 Announce Type: cross Abstract: AI coding tools are widely adopted, but most teams plateau at prompt-and-review without a framework for systematic progression. This paper presents th

The @aiDotEngineer Europe conference last week was a blast! Fun fact: @swyx & team pre-computed Gemini Embedding 2 vectors for all speakers …

Model ReleasesDGX agent

The @aiDotEngineer Europe conference last week was a blast! Fun fact: @swyx & team pre-computed Gemini Embedding 2 vectors for all speakers and sessions, and you can easily find similar sessions with

The entire team put a lot of effort into this benchmark. Document parsing and OCR certainly isn't solved, but now it is a bit easier to meas…

Model ReleasesDGX agent

The entire team put a lot of effort into this benchmark. Document parsing and OCR certainly isn't solved, but now it is a bit easier to measure 🦙 We’re open sourcing the first document OCR benchmark f

The model is not the agent. The harness is. You need to read this recent study, and a blog post from @hwchase17 ... (links below). It will r…

Model ReleasesDGX agent

The model is not the agent. The harness is. You need to read this recent study, and a blog post from @hwchase17 ... (links below). It will resonate deeply. This diagram from a recent paper captures so

The nextAI Solution to the NeurIPS 2023 LLM Efficiency Challenge

Model ReleasesDGX agent

arXiv:2604.09034v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) has significantly impacted the field of natural language processing, but their growing complexity ra

The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?

Model ReleasesDGX agent

arXiv:2601.07220v3 Announce Type: replace Abstract: Multilingual language models (LMs) promise broader NLP access, yet current systems deliver uneven performance across the world's languages. This sur

Thinking about trying Ollama Pro — how does it compare to Claude/Codex?

Model ReleasesDGX agent

This Reddit thread from r/ollama discusses user perspectives on **Ollama Pro** as a paid/upgraded tier compared to cloud-based AI coding assistants like Anthropic's Claude and OpenAI's Codex, likely f

TiAb Review Plugin: A Browser-Based Tool for AI-Assisted Title and Abstract Screening

Model ReleasesDGX agent

arXiv:2604.08602v1 Announce Type: cross Abstract: Background: Server-based screening tools impose subscription costs, while open-source alternatives require coding skills. Objectives: We developed a b

TIL @cognition usage has ~DOUBLED globally since these 2 launches. people are finding all sorts of creative usecases when u can compose agen…

Model ReleasesDGX agent

TIL @cognition usage has ~DOUBLED globally since these 2 launches. people are finding all sorts of creative usecases when u can compose agents together and make them proactive. agent recursion is all

TinyNeRV: Compact Neural Video Representations via Capacity Scaling, Distillation, and Low-Precision Inference

Model ReleasesDGX agent

arXiv:2604.09220v1 Announce Type: new Abstract: Implicit neural video representations encode entire video sequences within the parameters of a neural network and enable constant time frame reconstruct

← Previous
1…357358359360361…369
Next →