AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,778 results
Model Releases

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

DGX agent

arXiv:2606.25487v1 Announce Type: new Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an auto

model-releasesarxiv-cs-cl
25 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations

DGX agent

arXiv:2606.26041v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

How Small Can 6G Reason? Scaling Tiny-to-Small Language Models for AI-Native Networks

DGX agent

arXiv:2603.02156v2 Announce Type: replace-cross Abstract: Emerging 6G visions, reflected in ongoing standardization efforts within 3GPP, IETF, ETSI, ITU-T, and the O-RAN Alliance, increasingly charact

model-releasesarxiv-cs-ai
25 Jun 2026
Model Releases

https://openai.com/index/how-agents-are-transforming-work/

DGX agent

OpenAI discusses how AI agents are reshaping workplace productivity and operations, likely covering applications of autonomous AI systems in automating tasks, enhancing decision-making, and transformi

model-releasesopenai--x
25 Jun 2026
Model Releases

I spend 10 minutes a day trying a new app. http://matrix.build is my favorite app today. You can tell from a product whether the team's thin…

DGX agent

I spend 10 minutes a day trying a new app. http://matrix.build is my favorite app today. You can tell from a product whether the team's thinking is clear. The thing I love about matrix is it helps you

model-releaseszhipu-ai--x
25 Jun 2026
Model Releases

I'll be talking more about Claude Tag with @petergyang and at AIE with @_catwu. Let me know if there's anything you'd like us to dive into m…

DGX agent

I'll be talking more about Claude Tag with @petergyang and at AIE with @_catwu. Let me know if there's anything you'd like us to dive into more! Claude Tag is the next evolution of agents. It's a proa

model-releasesthariq--x
25 Jun 2026
Model Releases

Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation

DGX agent

arXiv:2411.15490v2 Announce Type: replace Abstract: Acute ischemic stroke (AIS) requires time-critical decision-making, where inaccurate interpretation of neuroimaging findings can lead to irreversibl

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Improving Zero-Shot Offline RL via Behavioral Task Sampling

DGX agent

arXiv:2604.25496v2 Announce Type: replace Abstract: Offline zero-shot reinforcement learning (RL) aims to learn agents that optimize unseen reward functions without additional environment interaction.

model-releasesarxiv-cs-ai
25 Jun 2026
Model Releases

In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Code…

DGX agent

In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Codex + GPT-5.5: - Judge score: 0.568 vs. 0.521 and 0.466 - Time

model-releasesfireworks-ai--x
25 Jun 2026
Model Releases

In-Context World Modeling for Robotic Control

DGX agent

arXiv:2606.26025v1 Announce Type: cross Abstract: Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

In HF GGUF section of models, we are emphasizing MTP heads with its own sign 𝗠𝗧𝗣

DGX agent

In HF GGUF section of models, we are emphasizing MTP heads with its own sign 𝗠𝗧𝗣 llama.cpp adds MTP for the Qwen3.6 family This is a significant milestone for the local AI ecosystem. The performance j

model-releasesgeorgi-gerganov--x
25 Jun 2026
Model Releases

In this episode, @OpenAI Chief Research Officer @markchen90 joins @allenpark to flambé shrimp, cook Korean stew, and chat about being at the…

DGX agent

In this episode, @OpenAI Chief Research Officer @markchen90 joins @allenpark to flambé shrimp, cook Korean stew, and chat about being at the frontier of AI research: why scaling laws and pre-training

model-releasesswyx--x
25 Jun 2026
Model Releases

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

DGX agent

arXiv:2606.19157v2 Announce Type: replace-cross Abstract: AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear wh

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

Internal Data Repetition Destroys Language Models

DGX agent

arXiv:2606.24998v1 Announce Type: new Abstract: Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition. Earlier cont

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Inverse Reinforcement Learning for Interpretable Keystroke Biomarkers in Parkinson's Disease

DGX agent

arXiv:2606.25270v1 Announce Type: new Abstract: Keystroke dynamics have been explored extensively as a passive digital biomarker for Parkinson's disease (PD), typically by extracting summary statistic

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

DGX agent

arXiv:2606.25984v1 Announce Type: cross Abstract: Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity

DGX agent

arXiv:2606.25343v1 Announce Type: new Abstract: Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significantly

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

DGX agent

arXiv:2606.25298v1 Announce Type: new Abstract: Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lackin

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Laplace--Fisher Gate Identities for Optimal Matrix-Gated Blended Score Estimation

DGX agent

arXiv:2606.25169v1 Announce Type: cross Abstract: Sampling from an unnormalized target by reversing an Ornstein--Uhlenbeck diffusion requires the score of each noise-perturbed marginal. Tweedie's iden

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Latent Block-Diffusion Temporal Point Processes: A Semi-Autoregressive Framework for Asynchronous Event Sequence Generation

DGX agent

arXiv:2606.24982v1 Announce Type: new Abstract: Modeling and sampling from the underlying distribution of asynchronous event sequences are crucial in various real-world applications, including social

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Learning Dynamical Systems from Multiple Sparse Datasets: A Hierarchical Bayesian Modeling Approach

DGX agent

arXiv:2606.24966v1 Announce Type: new Abstract: Estimating parameters of dynamical systems from sparse, noisy, and irregularly sampled data is often severely ill-conditioned. When multiple related dat

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Learning to Adapt: Reptile-D-Learning for Robust and Efficient Control Under Parametric Uncertainty

DGX agent

arXiv:2606.25659v1 Announce Type: new Abstract: Learning-based Lyapunov Control (LLC) provides formal stability guarantees for nonlinear systems, but its validity relies heavily on accurate system mod

model-releasesarxiv-cs-ro
25 Jun 2026
Model Releases

LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

DGX agent

arXiv:2606.25312v1 Announce Type: new Abstract: Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existin

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

LibEvoBench: Probing Temporal Knowledge Stratification in Code Generation Models

DGX agent

arXiv:2606.25402v1 Announce Type: cross Abstract: Large software projects often depend on older versions of libraries, even as APIs continue to evolve across releases. This creates a challenge for LLM

model-releasesarxiv-cs-ai
25 Jun 2026
Model Releases

LLM Performance on a Real, Double-Marked GCSE Benchmark

DGX agent

arXiv:2606.24973v1 Announce Type: new Abstract: We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

LLMs are getting better at writing GPU kernels. Multi-GPU kernels are the harder test. At @aiDotEngineer World's Fair, @simran_s_arora will …

DGX agent

LLMs are getting better at writing GPU kernels. Multi-GPU kernels are the harder test. At @aiDotEngineer World's Fair, @simran_s_arora will share ParallelKernelBench, an open-source benchmark built fr

model-releasestogether-ai--x
25 Jun 2026
Model Releases

MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios

DGX agent

arXiv:2606.24950v1 Announce Type: new Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamental

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Make Watergate Great Again.

DGX agent

Make Watergate Great Again. JD Vance: 'I think Nixon's historical legacy is enjoying a bit of a renaissance, and deservedly so. I joked that if Watergate happened tomorrow, it would be like a 12 hours

model-releasesanthropic--x
25 Jun 2026
Model Releases

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

DGX agent

arXiv:2604.05738v2 Announce Type: replace Abstract: Medical Vision-Language Models (Med-VLMs) have achieved expert-level proficiency in interpreting diagnostic imaging. However, current models are pre

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

Memory-Efficient Policy Libraries with Low-Rank Adaptation in Reinforcement Learning

DGX agent

arXiv:2606.25700v1 Announce Type: new Abstract: When fine-tuning Large Language Models (LLMs), there has been success in minimizing both memory usage and computation with Parameter-Efficient Fine-Tuni

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

MINIF2F-DAFNY: LLM-Guided Mathematical Theorem Proving via Auto-Active Verification

DGX agent

arXiv:2512.10187v3 Announce Type: replace Abstract: LLMs excel at reasoning, but validating their steps remains challenging. Formal verification offers a solution through mechanically checkable proofs

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Model-agnostic Mitigation Strategies of Data Imbalance for Regression

DGX agent

arXiv:2506.01486v2 Announce Type: replace Abstract: Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability.

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

DGX agent

arXiv:2606.26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But beh

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and Factorized Branch-and-Bound

DGX agent

arXiv:2606.25978v1 Announce Type: cross Abstract: Multi-agent goal recognition asks an observer to jointly infer which agents act together and what each team is trying to achieve, so the hypothesis sp

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Multi-Stream Temporal Fusion for Financial Fraud Detection

DGX agent

arXiv:2606.25007v1 Announce Type: new Abstract: Financial fraud detection in digital banking requires reasoning over multiple heterogeneous event streams -- transactions, login sessions, risk signals

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Multilingual Hematology Visual Question Answering Dataset

DGX agent

arXiv:2606.25246v1 Announce Type: cross Abstract: Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

DGX agent

arXiv:2606.26050v1 Announce Type: cross Abstract: Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ('Sue cried because'), it r

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

Ok, I am interviewing the legend @trq212 on Friday. I'm planning to ask him to show me: → His Claude Code setup and how he uses /goal and /l…

DGX agent

Ok, I am interviewing the legend @trq212 on Friday. I'm planning to ask him to show me: → His Claude Code setup and how he uses /goal and /loop and dynamic workflows → How he does planning with HTML +

model-releasesthariq--x
25 Jun 2026
Model Releases

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

DGX agent

arXiv:2606.25325v1 Announce Type: new Abstract: We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning tra

model-releasesarxiv-cs-ai
25 Jun 2026
Model Releases

On-Device Neural Architecture Search

DGX agent

arXiv:2606.24900v1 Announce Type: new Abstract: This paper proposes a new approach to near-sensor computing, in which a lightweight Neural Architecture Search (NAS) is performed directly on the deploy

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

OpenAI staggers GPT-5.6 rollout for government vetting, eyes 2027 IPO

DGX agent

OpenAI Group PBC will roll out its next model, GPT-5.6, to a small group of partners rather than the public at the request of the Trump administration, the latest sign that Washington now wants to rev

model-releasessiliconangle
25 Jun 2026
Model Releases

OpenAI will delay GPT-5.6 after Trump administration request

DGX agent

The Trump administration, apprehensive of potential security issues, has reportedly asked OpenAI to stagger the release of its next big-ticket model, GPT-5.6. The Information reported that OpenAI CEO

model-releasesthe-verge-ai
25 Jun 2026
Model Releases

Operator Boosting Produces Pareto-Efficient PDE Surrogates

DGX agent

arXiv:2606.17460v2 Announce Type: replace Abstract: Neural operators are widely used as surrogate solution maps for partial differential equations (PDEs), but full-size models can be costly to store,

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training

DGX agent

arXiv:2606.25906v1 Announce Type: new Abstract: With the advancement of artificial intelligence, research on oracle bone scripts has entered a new era. However, existing methods and benchmarks remain

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos

DGX agent

arXiv:2606.25245v1 Announce Type: new Abstract: Continuous 6-DoF pose estimation is essential for autonomous UAV operations. Yet, existing visual odometry and SLAM methods accumulate drift and yield o

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Our new AI policy is that the White House decides ad hoc, for whatever reasons it likes, who does and does not get access to frontier intell…

DGX agent

Our new AI policy is that the White House decides ad hoc, for whatever reasons it likes, who does and does not get access to frontier intelligence. This seems rather maximally terrible. The US Governm

model-releasesgary-marcus--x
25 Jun 2026
Model Releases

PatchINR: Patch-Based Implicit Neural Representations for Efficient and Scalable Inference

DGX agent

arXiv:2606.25534v1 Announce Type: new Abstract: Implicit Neural Representation (INR) provides an effective approach for continuous signal modeling, but classical per-pixel inference results in quadrat

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models

DGX agent

arXiv:2606.24952v1 Announce Type: new Abstract: A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should

model-releasesarxiv-cs-cl
25 Jun 2026
← Previous
1…170171172173174…475
Next →