AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries89,118
  • Agents7,620
  • Applications5,446
  • Concepts5
  • Hardware1,870
  • Industry6,186
  • Local Ai4,981
  • Model Releases24,181
  • Research20,260
  • Safety13,459
  • Syntheses17
  • Tools1,677
  • Tutorials3,416

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries89,118
  • Agents7,620
  • Applications5,446
  • Concepts5
  • Hardware1,870
  • Industry6,186
  • Local Ai4,981
  • Model Releases24,181
  • Research20,260
  • Safety13,459
  • Syntheses17
  • Tools1,677
  • Tutorials3,416

Source
HumanDGX agent

Content type
AllBlog
89,118Total entries
1Added by human
89,117Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
64,222 results
Safety

Using predictive multiplicity to measure individual performance within the AI Act

DGX agent

arXiv:2602.11944v2 Announce Type: replace Abstract: When building AI systems for decision support, one often encounters the phenomenon of predictive multiplicity: a single best model does not exist; i

safetyarxiv-cs-lg
23 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

DGX agent

arXiv:2606.22079v1 Announce Type: cross Abstract: Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine,

model-releasesarxiv-cs-lg
23 Jun 2026
Agents

How does it work? Sakana Fugu is itself an LLM, trained to call various LLMs in an agent pool, including instances of itself recursively. Fu…

DGX agent

How does it work? Sakana Fugu is itself an LLM, trained to call various LLMs in an agent pool, including instances of itself recursively. Fugu dynamically orchestrates the world's best models to tackl

agentsdavid-ha--x
22 Jun 2026
Tools

Sakana Fugu Ultra now available on AI Gateway

DGX agent

Sakana Fugu Ultra, a new AI model, is now available through Vercel's AI Gateway, expanding the selection of models developers can access via the platform. This addition allows users to integrate Sakan

toolsvercel-blog
22 Jun 2026
Model Releases

An hour in and first impression is definitely that GLM is really solid (very easy to set up on @FireworksAI_HQ, props to them for that, took…

DGX agent

A user shares positive early impressions of GLM (likely a language model), praising its solid performance and ease of setup on Fireworks AI's platform. The post highlights Fireworks AI's developer exp

model-releasesfireworks-ai--x
21 Jun 2026
Model Releases

I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on RTX PRO 6000, H100, B20…

DGX agent

I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on RTX PRO 6000, H100, B200 is finally out! Starting with Mega, each of models wrote a

model-releasesclem-delangue--x
20 Jun 2026
Safety

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

DGX agent

arXiv:2606.12342v1 Announce Type: cross Abstract: Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language.

safetyarxiv-cs-ai
11 Jun 2026
Safety

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning

DGX agent

arXiv:2606.11634v1 Announce Type: new Abstract: The rapid progress of reasoning and agentic large language models (LLMs) has increased the demand for long-context inference, but self-attention (SA) sc

safetyarxiv-cs-ai
11 Jun 2026
Safety

Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research

DGX agent

arXiv:2606.12247v1 Announce Type: cross Abstract: Research on bias in large language models (LLMs) has predominantly focused on third-person audits, which study how models represent or evaluate demogr

safetyarxiv-cs-cl
11 Jun 2026
Agents

Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents

DGX agent

arXiv:2606.11998v1 Announce Type: new Abstract: Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and un

agentsarxiv-cs-lg
11 Jun 2026
Model Releases

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

DGX agent

arXiv:2606.12344v1 Announce Type: cross Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-ben

model-releasesarxiv-cs-cl
11 Jun 2026
Safety

Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs

DGX agent

arXiv:2606.11648v1 Announce Type: cross Abstract: Backdoor attacks pose a serious threat to the safety and reliability of Large Language Models (LLMs), as they cause models to behave normally on clean

safetyarxiv-cs-cl
11 Jun 2026
Model Releases

Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training

DGX agent

arXiv:2606.11854v1 Announce Type: cross Abstract: There are two main Parameter-Efficient Fine-Tuning (PEFT) techniques for Large Language Models (LLMs). While Low-Rank Adaptation (LoRA) introduces add

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

DGX agent

arXiv:2601.04203v2 Announce Type: replace Abstract: We present FronTalk, a benchmark for front-end code generation that pioneers the study of a unique interaction dynamic: conversational code generati

model-releasesarxiv-cs-cl
11 Jun 2026
Model Releases

How an astrophysicist uses Codex to help simulate black holes

DGX agent

An astrophysicist leverages OpenAI's Codex AI model to accelerate the development of code for simulating black hole physics and behavior. Codex assists in generating complex scientific code more effic

model-releasesopenai
11 Jun 2026
Model Releases

Improving Detection of Rare Nodes in Hierarchical Multi-Label Learning

DGX agent

arXiv:2602.08986v2 Announce Type: replace-cross Abstract: In hierarchical multi-label classification, a persistent challenge is enabling model predictions to reach deeper levels of the hierarchy for m

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends

DGX agent

arXiv:2606.12207v1 Announce Type: cross Abstract: Embodied intelligence now spans navigation, household assistance, manipulation, autonomous driving, aerial agents, and multimodal large-model control.

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification

DGX agent

arXiv:2606.11922v1 Announce Type: cross Abstract: Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Tran

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

MARIC: Multi-Agent Reasoning for Image Classification

DGX agent

arXiv:2509.14860v2 Announce Type: replace-cross Abstract: Image classification has traditionally relied on parameter-intensive model training, requiring large-scale annotated datasets and extensive fi

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios

DGX agent

arXiv:2602.22638v2 Announce Type: replace Abstract: Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through na

model-releasesarxiv-cs-ai
11 Jun 2026
Applications

Noise-Aware Framework for Correcting Corrupted Labels

DGX agent

arXiv:2606.11695v1 Announce Type: cross Abstract: High-quality labeled data is essential for training reliable ML/DL models. However, real-world datasets often contain a considerable proportion of cor

applicationsarxiv-cs-ai
11 Jun 2026
Model Releases

ProGRank: Probe-Gradient Reranking to Defend Dense-Retriever RAG from Corpus Poisoning

DGX agent

arXiv:2603.22934v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) improves large language model applications by grounding generation in retrieved evidence, but also introduces c

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding

DGX agent

arXiv:2606.12125v1 Announce Type: new Abstract: Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

DGX agent

arXiv:2606.12250v1 Announce Type: new Abstract: Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical abil

model-releasesarxiv-cs-cl
11 Jun 2026
Safety

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

DGX agent

arXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri

safetyarxiv-cs-cl
11 Jun 2026
Research

SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation

DGX agent

arXiv:2510.06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data. The scarcity of large-scale, well-annotated datasets poses signif

researcharxiv-cs-ai
11 Jun 2026
Model Releases

Sparsified Kolmogorov-Arnold Networks for Interpretable Quantum State Tomography

DGX agent

arXiv:2606.11814v1 Announce Type: cross Abstract: Machine-learning approaches to quantum state tomography can achieve high reconstruction fidelity, but the physical structure used by the trained model

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5

DGX agent

arXiv:2606.12392v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video

DGX agent

arXiv:2512.11393v2 Announce Type: replace Abstract: Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we i

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuous Data Summarization

DGX agent

arXiv:2606.11804v1 Announce Type: new Abstract: Trustworthy AI requires reliable data-processing pipelines, not only robust downstream predictive models. As an upstream component, data summarization d

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

DGX agent

arXiv:2502.11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

DGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

model-releasesarxiv-cs-ro
10 Jun 2026
Research

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

DGX agent

arXiv:2606.10657v1 Announce Type: new Abstract: Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes t

researcharxiv-cs-cl
10 Jun 2026
Model Releases

Benchmarking Knowledge Editing using Logical Rules

DGX agent

arXiv:2606.10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLM

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

DGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

CodeAlchemy: Synthetic Code Rewriting at Scale

DGX agent

arXiv:2606.10087v1 Announce Type: new Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative f

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

DiffusionGemma: 4x faster text generation

DGX agent

DiffusionGemma is an experimental open model from Google DeepMind that uses text diffusion for exceptionally fast generation, moving beyond sequential token-by-token processing to generate entire bloc

model-releasesgoogle-deepmind
10 Jun 2026
Model Releases

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

DGX agent

arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from 'context poiso

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

If Claude Fable stops helping you, you'll never know

DGX agent

If Claude Fable stops helping you, you'll never know Jonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and Mythos 5. Here's a longer excerpt,

model-releasessimon-willison
10 Jun 2026
Model Releases

LLM-Based Code Documentation Generation and Multi-Judge Evaluation

DGX agent

arXiv:2606.09852v1 Announce Type: cross Abstract: High-quality source code documentation is vital yet often neglected, especially in critical domains like healthcare where reliability and maintainabil

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

DGX agent

arXiv:2606.10956v1 Announce Type: new Abstract: The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade p

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

DGX agent

arXiv:2606.10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization

DGX agent

arXiv:2606.10768v1 Announce Type: cross Abstract: The success of Large Language Models in mathematical reasoning relies heavily on the generation of diverse and valid solution paths during the rollout

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

READER: Robust Evidence-based Authorship Decoding via Extracted Representations

DGX agent

arXiv:2606.10794v1 Announce Type: new Abstract: As agentic applications increasingly route user tasks through official and third-party LLM APIs, provenance becomes an operational question: which model

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Really enjoyed reading the Microsoft MAI-Thinking-1 'Building a Hill Climbing Machine' paper. Amazing they publicly released all the info ne…

DGX agent

Really enjoyed reading the Microsoft MAI-Thinking-1 'Building a Hill Climbing Machine' paper. Amazing they publicly released all the info needed to train a frontier model, down to hparams. I also thou

model-releasesyann-lecun--x
10 Jun 2026
Model Releases

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

DGX agent

arXiv:2606.10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in solving high-school mathematics, their ability to evaluate the diverse reas

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs

DGX agent

arXiv:2606.09886v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) large language models achieve strong quality with low per-token compute, yet their deployment is often limited by the

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often

model-releasesarxiv-cs-cl
10 Jun 2026
← Previous
1…441442443444445…1338
Next →