AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,507 results
28 Jul 2026

Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age

Model ReleasesDGX agent

arXiv:2607.24341v1 Announce Type: new Abstract: Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal, or persona-b

SINT-Flow: Schema Integration using Large Language Model Workflows

Model ReleasesDGX agent

arXiv:2607.24492v1 Announce Type: new Abstract: The goal of schema integration is, given a set of input schemata or tables, to derive a global, unified schema that is able to represent the concepts, a

SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However,

SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion

Model ReleasesDGX agent

arXiv:2607.23580v1 Announce Type: new Abstract: Existing vector-sketch models treat recognition and generation as separate tasks, leaving a gap for streaming interfaces that must understand a drawing

Sling2Sim2Real: One-Shot Elastic System Identification for Non-Destructive Slingshot Policy Learning

Model ReleasesDGX agent

arXiv:2607.23268v1 Announce Type: new Abstract: Elastic object manipulation (EOM) involves highdimensional, nonlinear, and elastic deformations. The diverse deformation properties of elastic objects s

Small, Bias-Free, Blind and Convolutional Denoiser: A compact ConvNeXt U-Net for blind Gaussian color-image denoising

Model ReleasesDGX agent

arXiv:2607.22793v1 Announce Type: cross Abstract: We describe and evaluate BF-ConvUNeXt, a compact bias-free ConvNeXt U-Net for blind additive-white-Gaussian-noise color image denoising, combining fou

Smooth Learning with Hard Constraints via Legendre-Regularized Policies

Model ReleasesDGX agent

arXiv:2607.24007v1 Announce Type: cross Abstract: We revisit contextual optimization from the perspective of policy class design. A desirable policy class should be expressive enough to learn rich con

Source-Free Controlled Adaptation of Teachers for Continual Test-Time Adaptation

Model ReleasesDGX agent

arXiv:2607.23735v1 Announce Type: cross Abstract: In many real-world scenarios, encountering continual shifts in domain during inference is very common. Consequently, continual test-time adaptation (C

Sources: Moonshot is seeking access to more Nvidia Blackwell chips to prepare for Kimi K4's development, after training K3 on Nvidia chips, including Blackwell (The Information)

Model ReleasesDGX agent

The Information: Sources: Moonshot is seeking access to more Nvidia Blackwell chips to prepare for Kimi K4's development, after training K3 on Nvidia chips, including Blackwell — Beijing-based startup

Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.23474v1 Announce Type: new Abstract: This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

Model ReleasesDGX agent

arXiv:2607.22732v1 Announce Type: new Abstract: LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and

Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

Model ReleasesDGX agent

arXiv:2607.24701v1 Announce Type: new Abstract: Missing modalities in RGBT tracking often lead to incomplete and unstable multimodal feature representations that greatly degrade the performance. Exist

Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control

Model ReleasesDGX agent

arXiv:2607.10405v2 Announce Type: replace-cross Abstract: Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting users i

spec: add DSpark speculative decoding by wjinxu · Pull Request #25173 · ggml-org/llama.cpp

Model ReleasesDGX agent

It's time to experiment using DSpark! Please share your stats(pp/tg improvements). DSpark related stuff to check: DeepSpec - a deepseek-ai Collection DeepSeek-V4 with DSpark - DeepSeek-V4-Pro-DSpark &

Speed Reading Tool Powered by Artificial Intelligence for Students with ADHD, Dyslexia, and Short Attention Span

Model ReleasesDGX agent

arXiv:2307.14544v2 Announce Type: replace-cross Abstract: This paper presents an artificial intelligence tool designed to assist students with dyslexia, ADHD, and short attention spans in processing t

SPRKD: Effective Knowledge Distillation for Deep Neural Networks via Saddle Region Approximation

Model ReleasesDGX agent

arXiv:2607.23346v1 Announce Type: new Abstract: Modern deep neural networks are potent catalysts for scientific and industrial impact, yet excessive parameter counts impede deployment in low-compute s

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows

Model ReleasesDGX agent

arXiv:2607.23123v1 Announce Type: new Abstract: Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable deliverable produced

Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions

Model ReleasesDGX agent

arXiv:2603.20248v2 Announce Type: replace-cross Abstract: AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trust i

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

Model ReleasesDGX agent

arXiv:2607.22658v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain l

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

Model ReleasesDGX agent

arXiv:2607.24191v1 Announce Type: cross Abstract: Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key l

StAR: Segment Anything Reasoner

Model ReleasesDGX agent

arXiv:2603.14382v2 Announce Type: replace Abstract: As AI systems are being integrated more rapidly into diverse and complex real-world environments, the ability to perform holistic reasoning over an

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

Model ReleasesDGX agent

arXiv:2607.22798v1 Announce Type: cross Abstract: Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screen

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

Model ReleasesDGX agent

arXiv:2607.22708v1 Announce Type: new Abstract: Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the ac

Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls

Model ReleasesDGX agent

arXiv:2607.24519v1 Announce Type: cross Abstract: Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to negative con

Subject-Level Heterogeneity in EEG Motor Imagery Decoding: A Large-Scale Benchmark and Portfolio-Based Reduction of the Search Space

Model ReleasesDGX agent

arXiv:2607.22778v1 Announce Type: cross Abstract: Robust EEG motor imagery decoding remains limited by strong inter-individual variability, making it difficult to identify pipelines that generalize ac

Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.24054v1 Announce Type: new Abstract: A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer distinguishes i

SWE-rebench Multilingual Update (Go, Java, Python, Rust, TS). Evaluated: GLM-5.2, DeepSeek-V4 Pro, Qwen3.6-27B and others

Model ReleasesDGX agent

Hi everyone! We’ve just released a major update to the leaderboard! We are expanding beyond Python with a new multilingual slice featuring real-world software engineering tasks across 5 languages. Ope

SymStep: Symbolic Step Verification for Logical Reasoning

Model ReleasesDGX agent

arXiv:2607.23055v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps

Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

Model ReleasesDGX agent

arXiv:2607.23976v1 Announce Type: cross Abstract: Appending a two-word confirmation tag to a decision question -- 'Is X the better choice?' versus 'X is the better choice, right?' -- changes whether a

TEmBed-T: A Multi-Dimensional Benchmark for Table-Level Embeddings

Model ReleasesDGX agent

arXiv:2607.24130v1 Announce Type: cross Abstract: Tabular data is the dominant structured-data modality, and learning table representations has become a core research direction. Table-level embeddings

TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2

Model ReleasesDGX agent

arXiv:2606.19259v2 Announce Type: replace-cross Abstract: Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation model

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

Model ReleasesDGX agent

arXiv:2607.24063v1 Announce Type: new Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is tow

The Few-shot Dilemma: Over-prompting Large Language Models

Model ReleasesDGX agent

arXiv:2509.13196v2 Announce Type: replace Abstract: Over-prompting, a phenomenon where excessive examples in prompts lead to diminished performance in Large Language Models (LLMs), challenges the conv

The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.23335v1 Announce Type: new Abstract: Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that th

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

Model ReleasesDGX agent

arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked

The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts. 1. Yes, it looks r…

Model ReleasesDGX agent

The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts. 1. Yes, it looks relatively complicated, but it's essentially a scaled-up prod

The Label Complexity of Class-Conditional Coverage under Distribution Shift

Model ReleasesDGX agent

arXiv:2607.18088v2 Announce Type: replace-cross Abstract: Conformal prediction certifies that a classifier's prediction sets cover the truth, and that certificate is marginal. Many recognition benchma

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools,

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

Model ReleasesDGX agent

arXiv:2607.24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are tr

ThinkingCap-Qwen3.6-27B warrants a look

Model ReleasesDGX agent

It has only been two days since I move 100% from Qwen3.5-27B F16 to ThinkingCap-Qwen3.6-27B F16. Where I was getting tps in 30-40 range (depending on the size of the context), I am definitely getting

Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

Model ReleasesDGX agent

arXiv:2607.23054v1 Announce Type: cross Abstract: Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

Model ReleasesDGX agent

arXiv:2607.23951v1 Announce Type: new Abstract: Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods

Tines seeks to tame ‘wild code’ AI sprawl

Model ReleasesDGX agent

Workflow automation company Tines Security Services Ltd. today launched Tines 3B, an artificial intelligence-native platform for building, running and governing enterprise workflows, applications and

TLA^{+}-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation

Model ReleasesDGX agent

arXiv:2607.23425v1 Announce Type: cross Abstract: Large language models increasingly write TLA^{+} formal specifications from natural-language descriptions, but progress is hard to measure: existing r

TLRNet: Estimating Individual Treatment Effect based on Local Information and Single Learner Structure

Model ReleasesDGX agent

arXiv:2607.22762v1 Announce Type: cross Abstract: Causal inference has become a central issue across various fields, including computer science, statistics, economics, education, healthcare, and medic

Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations

Model ReleasesDGX agent

arXiv:2607.22610v1 Announce Type: new Abstract: When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies

TokenMem: Faithful Knowledge Injection for Frozen LLMs

Model ReleasesDGX agent

arXiv:2607.22625v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

Model ReleasesDGX agent

arXiv:2607.22954v1 Announce Type: new Abstract: Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world d

Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging

Model ReleasesDGX agent

arXiv:2607.24081v1 Announce Type: cross Abstract: Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challenges in devel

Transfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization

Model ReleasesDGX agent

arXiv:2607.23404v1 Announce Type: new Abstract: Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, exp

TRE: Training-Free Hallucination Detection for Diffusion Language Models

Model ReleasesDGX agent

arXiv:2607.22661v1 Announce Type: new Abstract: Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

Model ReleasesDGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query t

TriSP: Tri-Signal Structured Pruning for Large Language Models

Model ReleasesDGX agent

arXiv:2607.22587v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

Model ReleasesDGX agent

arXiv:2607.22727v1 Announce Type: new Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can

Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization

Model ReleasesDGX agent

arXiv:2607.23702v1 Announce Type: new Abstract: Manipulating objects with hidden internal state, such as a latched microwave, forces a robot to probe before it can act. Yet a robot that has solved an

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

Model ReleasesDGX agent

arXiv:2607.23458v1 Announce Type: new Abstract: Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-b

Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations

Model ReleasesDGX agent

arXiv:2607.23434v1 Announce Type: cross Abstract: Unexpected shocks recur in global operations, requiring decision rules that adapt as market and operating conditions change. Many operational systems

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

Model ReleasesDGX agent

arXiv:2508.00288v5 Announce Type: replace-cross Abstract: Aerial navigation is a fundamental yet underexplored capability in embodied intelligence, enabling agents to operate in large-scale, unstructu

Understanding Machine Unlearning Through the Lens of Mode Connectivity

Model ReleasesDGX agent

arXiv:2607.23970v1 Announce Type: cross Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss la

← Previous
1…6465666768…376
Next →