AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,428 results
11 Aug 2026

ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

ResearchDGX agent

arXiv:2608.08060v1 Announce Type: new Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pa

10 Aug 2026

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

Model ReleasesDGX agent

arXiv:2608.07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network anal

ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2608.06772v1 Announce Type: new Abstract: Accurate estimation of building energy use is essential for achieving carbon neutral and sustainable buildings. To better understand the influence of de

Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness

SafetyDGX agent

arXiv:2608.06613v1 Announce Type: cross Abstract: Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI artifacts remai

Flow-Corrected Shape Optimization: Taming Manifold Drift in High-Dimensional 3D Models

ResearchDGX agent

arXiv:2608.07199v1 Announce Type: new Abstract: Optimizing 3D shapes within the latent spaces of deep generative models is fundamental to computer assisted engineering, yet remains prone to a critical

Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders

ResearchDGX agent

arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promisi

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

SafetyDGX agent

arXiv:2608.06799v1 Announce Type: cross Abstract: Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

Model ReleasesDGX agent

arXiv:2608.06424v1 Announce Type: cross Abstract: Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utt

PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue

Model ReleasesDGX agent

arXiv:2608.06975v1 Announce Type: cross Abstract: Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on two fronts:

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

TutorialsDGX agent

arXiv:2608.07409v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent s

Very cool idea to have agents design complex systems by searching over the model structure itself. New research from Sakana AI introduces CE…

Model ReleasesDGX agent

Very cool idea to have agents design complex systems by searching over the model structure itself. New research from Sakana AI introduces CEDAR, which uses LLM agents to write, simulate, and refine sy

We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work…

Model ReleasesDGX agent

We’re expanding our cybersecurity initiative Daybreak and introducing GPT-5.6-Cyber, a new model for advanced, authorized cybersecurity work. As the threat landscape evolves, we’re putting frontier in

ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?

Local AiDGX agent

arXiv:2608.07033v1 Announce Type: new Abstract: This work investigates whether Electroencephalograph (EEG) foundation models (EFMs) can be made faster and locally deployable without sacrificing accura

8 Aug 2026

OpenAI reveals upcoming Astra model may possess ‘critical’ hacking capabilities

IndustryDGX agent

OpenAI Group PBC today disclosed that one of its unreleased large language models may pose a significant cybersecurity risk. The algorithm, which is known as Astra, was first detailed last week. OpenA

7 Aug 2026

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

Model ReleasesDGX agent

arXiv:2608.06243v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcom

DynaPix: Can Vision-Language Models Identify the Exact Future?

Model ReleasesDGX agent

arXiv:2608.05505v1 Announce Type: new Abstract: Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking ima

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

ResearchDGX agent

arXiv:2608.05265v1 Announce Type: cross Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in r

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

SafetyDGX agent

arXiv:2608.06332v1 Announce Type: new Abstract: Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning a

Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations

Model ReleasesDGX agent

arXiv:2604.17359v2 Announce Type: replace-cross Abstract: Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no real

6 Aug 2026

An accurate characterization of the arc of AI is that it is shaped by two trends: 1. Moving more and more logic to a neural model for tasks …

AgentsDGX agent

An accurate characterization of the arc of AI is that it is shaped by two trends: 1. Moving more and more logic to a neural model for tasks where training data can be densely sampled (e.g. the shift f

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

SafetyDGX agent

arXiv:2608.04302v1 Announce Type: new Abstract: Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate acc

Faster-WAM: Efficient Inference-Time Future Conditioning for Robust World Action Models

Model ReleasesDGX agent

arXiv:2608.04404v1 Announce Type: new Abstract: World Action Models (WAMs) improve robot manipulation by learning how the environment evolves beyond the current observation. However, existing approach

From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models

Model ReleasesDGX agent

arXiv:2608.04200v1 Announce Type: cross Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models

Model ReleasesDGX agent

arXiv:2608.04351v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognitio

i just spent weeks rewriting my webUI from scratch, getting rid of all AI slop within the codebase and switching it over to a proper lightweight framework (alpine.js). i am now comfortable suggesting it as an alternative to openwebUI, librechat and the like! it is made for local models

Local AiDGX agent

[Fully open source under GPL3, made from the ground up for use with local models, no subscriptions, no corporate backing] When i first started this, it was meant to be a fully lightweight, extremely m

LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling

Model ReleasesDGX agent

arXiv:2608.04147v1 Announce Type: cross Abstract: Label noise is common in medical imaging datasets due to factors such as inter-rater variability, annotation errors, and ambiguous cases. This can sev

Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling

ResearchDGX agent

arXiv:2608.02347v2 Announce Type: replace Abstract: Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fi

Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?

Model ReleasesDGX agent

arXiv:2608.05097v1 Announce Type: new Abstract: Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same

Teaching Foundation Models to Read mmWave: Pose-Guided Kinematic Representation for Human Behavior Understanding

Model ReleasesDGX agent

arXiv:2608.04127v1 Announce Type: new Abstract: Large language model agents need to perceive human behavior in physical environments. Millimeter-wave (mmWave) radar provides a privacy-friendly and con

5 Aug 2026

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

SafetyDGX agent

arXiv:2608.02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in

A game theory for foundation models shows new paths to rational cooperation through similarity inference

SafetyDGX agent

arXiv:2608.03958v1 Announce Type: new Abstract: As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing t

Can Text-to-Image Models Draw from the Right Frame of Reference?

Model ReleasesDGX agent

arXiv:2608.03357v1 Announce Type: new Abstract: Spatial instruction following has become a crucial requirement for text-to-image (T2I) generation. A common challenge arises when directional expression

Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models

ResearchDGX agent

arXiv:2608.03160v1 Announce Type: cross Abstract: When asked which of two events came first, video large language models can fail in two opposite ways: cave to a false claim, or reject a true one. Pri

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

Model ReleasesDGX agent

arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Model ReleasesDGX agent

arXiv:2608.02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language pr

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

Model ReleasesDGX agent

arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a ho

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

HardwareDGX agent

arXiv:2608.03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticip

MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models

ResearchDGX agent

arXiv:2608.03769v1 Announce Type: cross Abstract: Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundament

MIDI-LLM: Improving Text-to-MIDI Music Generation via Adapting Large Language Models

Model ReleasesDGX agent

arXiv:2511.03942v2 Announce Type: replace-cross Abstract: We present MIDI-LLM, a recipe that improves multitrack text-to-MIDI generation via adapting Large Language Models (LLMs). MIDI-LLM expands an

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Model ReleasesDGX agent

arXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to f

Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models

SafetyDGX agent

arXiv:2511.21214v4 Announce Type: replace-cross Abstract: Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We stu

SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay

Model ReleasesDGX agent

arXiv:2608.03063v1 Announce Type: new Abstract: Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.03580v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

SafetyDGX agent

arXiv:2608.02951v1 Announce Type: cross Abstract: Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-free methods

4 Aug 2026

A Heuristic Perspective on Debiasing Language Models

SafetyDGX agent

arXiv:2608.00622v1 Announce Type: new Abstract: Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm. Existing m

Bole: Efficient Tree Speculation for Hybrid-Attention Language Models

HardwareDGX agent

arXiv:2608.01651v1 Announce Type: cross Abstract: Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autore

Certifying Plans under Model Mismatch: A Trilemma for Reachability from Scarce Data

Model ReleasesDGX agent

arXiv:2608.02453v1 Announce Type: new Abstract: Sim-to-real policies are designed under nominal dynamics, but target-system trials may yield only a few isolated one-step transitions. We study pre-exec

Cluster-Aware Over-the-Air Federated Learning with Energy-Harvesting Devices: From Global Training to Model Personalization

Model ReleasesDGX agent

arXiv:2608.01426v1 Announce Type: new Abstract: Federated learning (FL) enables distributed optimization and learning across decentralized edge devices while preserving data privacy, but its performan

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.01644v1 Announce Type: new Abstract: In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the pre

I benchmarked the 4 models I had pulled. The 1.1GB one beat the 2GB one at math and lost badly at extraction.

Model ReleasesDGX agent

152 generations, deterministic grading (exact number/string/JSON/regex), no LLM judge on my 16GB laptop. task type | deepseek-r1:1.5b (1.1GB) | llama3.2:3b (2.0GB) | gemma:2b | codellama 7b arithmetic

Learning the Pareto Frontier of Predictive Models under Distribution Shift

ApplicationsDGX agent

arXiv:2608.00632v1 Announce Type: new Abstract: Modern machine learning pipelines increasingly rely on reusing pretrained and foundation models across downstream tasks. These pretrained models can dif

Look Where It Matters: Adaptive Visual Refinement for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.02197v1 Announce Type: new Abstract: Visual representations of VLA models remain unreliable for spatially precise robotic manipulation. We uncover that vision encoders in VLAs also exhibit

MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models

SafetyDGX agent

arXiv:2608.01012v1 Announce Type: new Abstract: Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagno

Mistral releases Shieldstral, a 3B multimodal safety classifier that it says matches models up to 7x its size on text safety, available under Apache 2.0 (Mistral AI Blog)

Model ReleasesDGX agent

Mistral AI Blog: Mistral releases Shieldstral, a 3B multimodal safety classifier that it says matches models up to 7x its size on text safety, available under Apache 2.0 — Every product that ships a m

MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling

SafetyDGX agent

arXiv:2604.02817v2 Announce Type: replace Abstract: Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel

Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelligence from Raw Earth Observation Streams

Model ReleasesDGX agent

arXiv:2608.00012v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster

PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs

SafetyDGX agent

arXiv:2608.01458v1 Announce Type: new Abstract: Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the s

Progressive^2: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression

ResearchDGX agent

arXiv:2608.00129v1 Announce Type: new Abstract: Knowledge distillation (KD) is a widely utilized technique for transferring knowledge from a large model (the teacher) to a smaller model (the student).

Releasing Pokee-Isaac 28B — the world’s first real 10M-token context frontier-class agentic model, deployable on a single GPU (starting from…

Model ReleasesDGX agent

Releasing Pokee-Isaac 28B — the world’s first real 10M-token context frontier-class agentic model, deployable on a single GPU (starting from RTX 4090 or equivalent). New proprietary non-decoder-only a

Role Steering of Language Models for Social Simulations

Model ReleasesDGX agent

arXiv:2608.00023v1 Announce Type: new Abstract: Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated populat

← Previous
1…7778798081…1008
Next →