AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
3 Jul 2026

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

Model ReleasesDGX agent

arXiv:2607.01751v1 Announce Type: cross Abstract: Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right ti

Meta-Benchmarks for Financial-Services LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.01740v1 Announce Type: new Abstract: Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model th

Meta to release new AI model with advanced coding capabilities ‘soon’

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Meta Platforms Inc. is gearing up to release a new version of its flagship Muse Spark artificial intelligence model. Alexandr Wang, the company’s chief AI officer, wrote on X today that the update wil

MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2506.09105v3 Announce Type: replace-cross Abstract: We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter-ef

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

Model ReleasesDGX agent

arXiv:2607.01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes va

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

Model ReleasesDGX agent

arXiv:2607.01627v1 Announce Type: cross Abstract: Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development. A difficul

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

Model ReleasesDGX agent

arXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

Model ReleasesDGX agent

arXiv:2607.01814v1 Announce Type: new Abstract: Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and reproducibility. T

Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space

Model ReleasesDGX agent

arXiv:2607.01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.~Most existing appr

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

Model ReleasesDGX agent

arXiv:2607.01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

Model ReleasesDGX agent

arXiv:2604.04532v2 Announce Type: replace-cross Abstract: Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language

mupscaling small models: Principled warm starts and hyperparameter transfer

Model ReleasesDGX agent

arXiv:2602.10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve effic

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

Model ReleasesDGX agent

arXiv:2601.01095v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand tempo

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

Model ReleasesDGX agent

arXiv:2607.01378v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world depl

NEUROSYMLAND: Neuro-Symbolic Landing-Site Assessment for Robust and Edge-Deployable UAV Autonomy

Model ReleasesDGX agent

arXiv:2607.02277v1 Announce Type: new Abstract: Safe landing-site assessment in unstructured environments remains a key challenge for autonomous UAV deployment, as vision-only learning approaches ofte

Office Comprehension Benchmark

Model ReleasesDGX agent

arXiv:2607.01245v1 Announce Type: cross Abstract: We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension

OmniGAIA: Towards Native Omni-Modal AI Agents

Model ReleasesDGX agent

arXiv:2602.22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to i

On the Limits of Steering Vectors for Preference-Aligned Generation

Model ReleasesDGX agent

arXiv:2607.01802v1 Announce Type: new Abstract: Steering vectors have emerged as a promising approach to controlled text generation, offering interpretable, training-free mechanisms for shaping model

On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain

Model ReleasesDGX agent

arXiv:2607.01444v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models offer inference speedups via selective activation but impose substantial memory requirements because the whole network

One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective

Model ReleasesDGX agent

arXiv:2607.02292v1 Announce Type: new Abstract: Neural quantum states (NQS) provide a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, au

Open Source AI Gap Map

Model ReleasesDGX agent

Open Source AI Gap Map Current AI is 'a global partnership building a public option for AI', founded as a non-profit at the AI Action Summit in Paris in February 2025 and backed by serious capital ($4

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

Model ReleasesDGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

Model ReleasesDGX agent

arXiv:2607.01531v1 Announce Type: new Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networ

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

Model ReleasesDGX agent

arXiv:2607.02461v1 Announce Type: cross Abstract: Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make infe

PACE: A Proxy for Agentic Capability Evaluation

Model ReleasesDGX agent

arXiv:2607.02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation c

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation

Model ReleasesDGX agent

arXiv:2607.01883v1 Announce Type: new Abstract: Code is the medium through which large language models generate structured artifacts: charts, scientific figures, vector graphics, CAD models, 3D scenes

Parameter Golf: What Really Works?

Model ReleasesDGX agent

arXiv:2607.01517v1 Announce Type: new Abstract: How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which particip

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2607.01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribu

Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

Model ReleasesDGX agent

arXiv:2506.12311v4 Announce Type: replace Abstract: Text-to-speech (TTS) for Modern Hebrew is challenged by the language's orthographic complexity, with existing solutions ignoring underspecified phon

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

Model ReleasesDGX agent

arXiv:2607.01938v1 Announce Type: cross Abstract: Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action

Population-Scale Segmentation of Penile Tissue in DIXON MRI using Deep Learning for Quantitative Phenotyping in Male Reproductive Health

Model ReleasesDGX agent

arXiv:2607.02127v1 Announce Type: cross Abstract: Penile measurement is clinically relevant across male reproductive and urogenital health, including conditions such as micropenis, congenital and endo

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

Model ReleasesDGX agent

arXiv:2606.20950v2 Announce Type: replace Abstract: Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way

PPTArena: A Benchmark for PowerPoint Editing

Model ReleasesDGX agent

arXiv:2512.03042v3 Announce Type: replace-cross Abstract: We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions. Unl

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

Model ReleasesDGX agent

arXiv:2607.01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing a

PreScience: A Dataset and Benchmark for Scientific Forecasting

Model ReleasesDGX agent

arXiv:2602.20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark fo

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Model ReleasesDGX agent

arXiv:2607.02512v1 Announce Type: cross Abstract: Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking

Prompt Framing Distorts Count-Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Model ReleasesDGX agent

arXiv:2607.01240v1 Announce Type: cross Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding i

Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

Model ReleasesDGX agent

arXiv:2510.04391v5 Announce Type: replace Abstract: Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic large la

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Model ReleasesDGX agent

arXiv:2510.04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactio

QFedAgent: Quantum-Enhanced Personalized Federated Learning for Multi-Agent Activity Recognition

Model ReleasesDGX agent

arXiv:2607.02426v1 Announce Type: cross Abstract: Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, making it suitable for privacy-sensi

RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.02185v1 Announce Type: cross Abstract: Deep learning has achieved remarkable performance in medical image segmentation, yet it suffers from critical limitations: mathematical intractability

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Model ReleasesDGX agent

arXiv:2607.02504v1 Announce Type: cross Abstract: Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on extbf{sp

Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration

Model ReleasesDGX agent

arXiv:2603.06001v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasing

Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

Model ReleasesDGX agent

arXiv:2607.01951v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat

Rocket Report: Indian startup nears first launch; SpaceX's millenary milestone

Model ReleasesDGX agent

India's first space unicorn is preparing for its first orbital launch , with Skyroot Aerospace developing the Vikram series of launch vehicles . In 2025, SpaceX posted a net loss of $4.9 billion , hig

RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

Model ReleasesDGX agent

arXiv:2607.01388v1 Announce Type: new Abstract: Multi-step symbolic reasoning is essential for robust financial analysis, yet most benchmarks neglect intermediate reasoning steps. FINCHAIN introduced

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.01876v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and cross-modal

Safety Targeted Embedding Exploit via Refinement

Model ReleasesDGX agent

arXiv:2607.01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-r

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

Model ReleasesDGX agent

arXiv:2607.01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testin

Scaling Trends for Lie Detector Oversight in Preference Learning

Model ReleasesDGX agent

arXiv:2607.01567v1 Announce Type: new Abstract: Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave,

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

Model ReleasesDGX agent

arXiv:2607.01612v1 Announce Type: new Abstract: Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering

SCAPE: Accurate and Efficient LLM Training with Extreme Sparse Communication

Model ReleasesDGX agent

arXiv:2607.01678v1 Announce Type: new Abstract: Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, w

Self-Gating Attention for Efficient Time Series Forecasting

Model ReleasesDGX agent

arXiv:2607.02344v1 Announce Type: cross Abstract: Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture temporal d

Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment

Model ReleasesDGX agent

arXiv:2607.01674v1 Announce Type: new Abstract: In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freezing a pretrai

SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation

Model ReleasesDGX agent

arXiv:2607.01766v1 Announce Type: new Abstract: LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static output. Dynamic

Some notes from @aiDotEngineer world fair: > the energy was incredible. it's magical to have a large group of smart, hungry, technical, and …

Model ReleasesDGX agent

Some notes from @aiDotEngineer world fair: > the energy was incredible. it's magical to have a large group of smart, hungry, technical, and driven people learning from each other under one roof. > the

Something about this year’s @aiDotEngineer World’s Fair just hit different. Last year was the year of “let the agents rip.” This year was th…

Model ReleasesDGX agent

Something about this year’s @aiDotEngineer World’s Fair just hit different. Last year was the year of “let the agents rip.” This year was the year of realizing that autonomy without structure creates

Sources: Alibaba has banned employees from using Claude Code and asked them to remove all Claude models from their work computers, citing security concerns (The Information)

Model ReleasesDGX agent

The Information: Sources: Alibaba has banned employees from using Claude Code and asked them to remove all Claude models from their work computers, citing security concerns — Alibaba Group has banned

Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters

Model ReleasesDGX agent

arXiv:2607.01893v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only t

Spectral Imbalance Causes Forgetting in Low-Rank Continual Adaptation

Model ReleasesDGX agent

arXiv:2602.00722v2 Announce Type: replace Abstract: Parameter-efficient continual learning aims to adapt pre-trained models to sequential tasks without forgetting previously acquired knowledge. Most e

← Previous
1…107108109110111…377
Next →