AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,577 results
Model Releases

LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models

DGX agent

arXiv:2504.02327v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive acce

model-releasesarxiv-cs-cl
3 Jul 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

DGX agent

arXiv:2607.02466v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions,

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens

DGX agent

arXiv:2507.02964v2 Announce Type: replace-cross Abstract: The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scalable

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Liquid Latent State Dynamics for Interpretable Turbofan Degradation Modeling

DGX agent

arXiv:2607.01986v1 Announce Type: new Abstract: Multivariate time-series models for prognostics are often evaluated by point prediction accuracy, yet their internal states rarely expose a coherent deg

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability

DGX agent

arXiv:2607.01247v1 Announce Type: cross Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate s

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Locality-Aware Continual Unlearning for Diffusion Models

DGX agent

arXiv:2512.02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations ar

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

DGX agent

arXiv:2607.01764v1 Announce Type: new Abstract: Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar tha

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation

DGX agent

arXiv:2508.16674v2 Announce Type: replace-cross Abstract: Medical report understanding from real-world document images is essential for generating patient-facing explanations and enabling structured i

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

DGX agent

arXiv:2607.01751v1 Announce Type: cross Abstract: Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right ti

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Meta-Benchmarks for Financial-Services LLM Evaluation

DGX agent

arXiv:2607.01740v1 Announce Type: new Abstract: Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model th

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Meta to release new AI model with advanced coding capabilities ‘soon’

DGX agent

Meta Platforms Inc. is gearing up to release a new version of its flagship Muse Spark artificial intelligence model. Alexandr Wang, the company’s chief AI officer, wrote on X today that the update wil

model-releasessiliconangle
3 Jul 2026
Model Releases

MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2506.09105v3 Announce Type: replace-cross Abstract: We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter-ef

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

DGX agent

arXiv:2607.01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes va

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

DGX agent

arXiv:2607.01627v1 Announce Type: cross Abstract: Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development. A difficul

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

DGX agent

arXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

DGX agent

arXiv:2607.01814v1 Announce Type: new Abstract: Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and reproducibility. T

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space

DGX agent

arXiv:2607.01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.~Most existing appr

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

DGX agent

arXiv:2607.01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

DGX agent

arXiv:2604.04532v2 Announce Type: replace-cross Abstract: Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

mupscaling small models: Principled warm starts and hyperparameter transfer

DGX agent

arXiv:2602.10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve effic

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

DGX agent

arXiv:2601.01095v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand tempo

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

DGX agent

arXiv:2607.01378v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world depl

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

NEUROSYMLAND: Neuro-Symbolic Landing-Site Assessment for Robust and Edge-Deployable UAV Autonomy

DGX agent

arXiv:2607.02277v1 Announce Type: new Abstract: Safe landing-site assessment in unstructured environments remains a key challenge for autonomous UAV deployment, as vision-only learning approaches ofte

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

Office Comprehension Benchmark

DGX agent

arXiv:2607.01245v1 Announce Type: cross Abstract: We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OmniGAIA: Towards Native Omni-Modal AI Agents

DGX agent

arXiv:2602.22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to i

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

On the Limits of Steering Vectors for Preference-Aligned Generation

DGX agent

arXiv:2607.01802v1 Announce Type: new Abstract: Steering vectors have emerged as a promising approach to controlled text generation, offering interpretable, training-free mechanisms for shaping model

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain

DGX agent

arXiv:2607.01444v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models offer inference speedups via selective activation but impose substantial memory requirements because the whole network

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective

DGX agent

arXiv:2607.02292v1 Announce Type: new Abstract: Neural quantum states (NQS) provide a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, au

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Open Source AI Gap Map

DGX agent

Open Source AI Gap Map Current AI is 'a global partnership building a public option for AI', founded as a non-profit at the AI Action Summit in Paris in February 2025 and backed by serious capital ($4

model-releasessimon-willison
3 Jul 2026
Model Releases

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

DGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

DGX agent

arXiv:2607.01531v1 Announce Type: new Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networ

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

DGX agent

arXiv:2607.02461v1 Announce Type: cross Abstract: Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make infe

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

PACE: A Proxy for Agentic Capability Evaluation

DGX agent

arXiv:2607.02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation c

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation

DGX agent

arXiv:2607.01883v1 Announce Type: new Abstract: Code is the medium through which large language models generate structured artifacts: charts, scientific figures, vector graphics, CAD models, 3D scenes

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Parameter Golf: What Really Works?

DGX agent

arXiv:2607.01517v1 Announce Type: new Abstract: How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which particip

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

DGX agent

arXiv:2607.01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribu

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

DGX agent

arXiv:2506.12311v4 Announce Type: replace Abstract: Text-to-speech (TTS) for Modern Hebrew is challenged by the language's orthographic complexity, with existing solutions ignoring underspecified phon

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

DGX agent

arXiv:2607.01938v1 Announce Type: cross Abstract: Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Population-Scale Segmentation of Penile Tissue in DIXON MRI using Deep Learning for Quantitative Phenotyping in Male Reproductive Health

DGX agent

arXiv:2607.02127v1 Announce Type: cross Abstract: Penile measurement is clinically relevant across male reproductive and urogenital health, including conditions such as micropenis, congenital and endo

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

DGX agent

arXiv:2606.20950v2 Announce Type: replace Abstract: Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

PPTArena: A Benchmark for PowerPoint Editing

DGX agent

arXiv:2512.03042v3 Announce Type: replace-cross Abstract: We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions. Unl

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

DGX agent

arXiv:2607.01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing a

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

PreScience: A Dataset and Benchmark for Scientific Forecasting

DGX agent

arXiv:2602.20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark fo

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

DGX agent

arXiv:2607.02512v1 Announce Type: cross Abstract: Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Prompt Framing Distorts Count-Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

DGX agent

arXiv:2607.01240v1 Announce Type: cross Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding i

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate

DGX agent

arXiv:2510.04391v5 Announce Type: replace Abstract: Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic large la

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

DGX agent

arXiv:2510.04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactio

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

QFedAgent: Quantum-Enhanced Personalized Federated Learning for Multi-Agent Activity Recognition

DGX agent

arXiv:2607.02426v1 Announce Type: cross Abstract: Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, making it suitable for privacy-sensi

model-releasesarxiv-cs-ai
3 Jul 2026
← Previous
1…133134135136137…471
Next →