AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,223
  • Agents7,699
  • Applications5,506
  • Concepts5
  • Hardware1,889
  • Industry6,186
  • Local Ai5,045
  • Model Releases24,499
  • Research20,615
  • Safety13,633
  • Syntheses17
  • Tools1,677
  • Tutorials3,452

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,223
  • Agents7,699
  • Applications5,506
  • Concepts5
  • Hardware1,889
  • Industry6,186
  • Local Ai5,045
  • Model Releases24,499
  • Research20,615
  • Safety13,633
  • Syntheses17
  • Tools1,677
  • Tutorials3,452

Source
HumanDGX agent

Content type
AllBlog
90,223Total entries
1Added by human
90,222Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,101 results
Model Releases

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

DGX agent

arXiv:2606.12250v1 Announce Type: new Abstract: Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical abil

model-releasesarxiv-cs-cl
11 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

DGX agent

arXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri

safetyarxiv-cs-cl
11 Jun 2026
Research

SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation

DGX agent

arXiv:2510.06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data. The scarcity of large-scale, well-annotated datasets poses signif

researcharxiv-cs-ai
11 Jun 2026
Model Releases

Sparsified Kolmogorov-Arnold Networks for Interpretable Quantum State Tomography

DGX agent

arXiv:2606.11814v1 Announce Type: cross Abstract: Machine-learning approaches to quantum state tomography can achieve high reconstruction fidelity, but the physical structure used by the trained model

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5

DGX agent

arXiv:2606.12392v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video

DGX agent

arXiv:2512.11393v2 Announce Type: replace Abstract: Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we i

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuous Data Summarization

DGX agent

arXiv:2606.11804v1 Announce Type: new Abstract: Trustworthy AI requires reliable data-processing pipelines, not only robust downstream predictive models. As an upstream component, data summarization d

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

DGX agent

arXiv:2502.11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

DGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

model-releasesarxiv-cs-ro
10 Jun 2026
Research

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

DGX agent

arXiv:2606.10657v1 Announce Type: new Abstract: Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes t

researcharxiv-cs-cl
10 Jun 2026
Model Releases

Benchmarking Knowledge Editing using Logical Rules

DGX agent

arXiv:2606.10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLM

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

DGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

CodeAlchemy: Synthetic Code Rewriting at Scale

DGX agent

arXiv:2606.10087v1 Announce Type: new Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative f

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

DiffusionGemma: 4x faster text generation

DGX agent

DiffusionGemma is an experimental open model from Google DeepMind that uses text diffusion for exceptionally fast generation, moving beyond sequential token-by-token processing to generate entire bloc

model-releasesgoogle-deepmind
10 Jun 2026
Model Releases

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

DGX agent

arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from 'context poiso

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

If Claude Fable stops helping you, you'll never know

DGX agent

If Claude Fable stops helping you, you'll never know Jonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and Mythos 5. Here's a longer excerpt,

model-releasessimon-willison
10 Jun 2026
Model Releases

LLM-Based Code Documentation Generation and Multi-Judge Evaluation

DGX agent

arXiv:2606.09852v1 Announce Type: cross Abstract: High-quality source code documentation is vital yet often neglected, especially in critical domains like healthcare where reliability and maintainabil

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

DGX agent

arXiv:2606.10956v1 Announce Type: new Abstract: The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade p

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

DGX agent

arXiv:2606.10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization

DGX agent

arXiv:2606.10768v1 Announce Type: cross Abstract: The success of Large Language Models in mathematical reasoning relies heavily on the generation of diverse and valid solution paths during the rollout

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

READER: Robust Evidence-based Authorship Decoding via Extracted Representations

DGX agent

arXiv:2606.10794v1 Announce Type: new Abstract: As agentic applications increasingly route user tasks through official and third-party LLM APIs, provenance becomes an operational question: which model

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Really enjoyed reading the Microsoft MAI-Thinking-1 'Building a Hill Climbing Machine' paper. Amazing they publicly released all the info ne…

DGX agent

Really enjoyed reading the Microsoft MAI-Thinking-1 'Building a Hill Climbing Machine' paper. Amazing they publicly released all the info needed to train a frontier model, down to hparams. I also thou

model-releasesyann-lecun--x
10 Jun 2026
Model Releases

RealMath-Eval: Why SOTA Judges Struggle with Real Human Reasoning

DGX agent

arXiv:2606.10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in solving high-school mathematics, their ability to evaluate the diverse reas

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs

DGX agent

arXiv:2606.09886v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (MoE) large language models achieve strong quality with low per-token compute, yet their deployment is often limited by the

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often

model-releasesarxiv-cs-cl
10 Jun 2026
Model Releases

STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios

DGX agent

arXiv:2606.10394v1 Announce Type: new Abstract: Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existin

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

DGX agent

arXiv:2606.11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However,

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

The Order Matters: Sequential Fine-Tuning of LLaMA for Coherent Automated Essay Scoring

DGX agent

arXiv:2606.10327v1 Announce Type: new Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.g., lead, claim, evidence, conclusion), yet most approaches treat

model-releasesarxiv-cs-cl
10 Jun 2026
Research

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

DGX agent

arXiv:2606.10278v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals. While recent advances in deep learning have signific

researcharxiv-cs-ai
10 Jun 2026
Model Releases

TRAPS: Therapeutic Response Analysis via Pathway-informed Stratification

DGX agent

arXiv:2606.09898v1 Announce Type: new Abstract: Cancer treatment planning requires decisions across multiple clinical dimensions at once. Clinicians must determine whether a patient should receive tar

model-releasesarxiv-cs-lg
10 Jun 2026
Model Releases

wooh https://x.com/shadcn/status/2064671802509410806?s=46

DGX agent

wooh https://x.com/shadcn/status/2064671802509410806?s=46 You have Claude Fable for only a few days. Here's how to make the most of it. Introducing /improve: use your most capable model to audit your

model-releasesswyx--x
10 Jun 2026
Model Releases

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

DGX agent

arXiv:2606.08669v1 Announce Type: cross Abstract: Voice biometric systems face growing threats from spoofing attacks, yet the evaluation of detection models remains inconsistent across datasets. To in

model-releasesarxiv-cs-lg
9 Jun 2026
Safety

Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks

DGX agent

arXiv:2604.01039v2 Announce Type: replace-cross Abstract: System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensitive

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems

DGX agent

arXiv:2606.07805v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) from passive assistants to autonomous, execution-capable agents has introduced critical operational

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP

DGX agent

arXiv:2505.11189v3 Announce Type: replace Abstract: Large language models (LLMs) can amplify misinformation, undermining societal goals such as the UN SDGs. We study three documented drivers of misinf

model-releasesarxiv-cs-ai
9 Jun 2026
Research

Capacity, Not Format: Rethinking Structured Reasoning Failures

DGX agent

arXiv:2606.09410v1 Announce Type: new Abstract: Prior work treats structured output as a reasoning tax, but this framing is incomplete: the cost of formatting depends strongly on a model's spare capac

researcharxiv-cs-ai
9 Jun 2026
Model Releases

CATPO: Critique-Augmented Tree Policy Optimization

DGX agent

arXiv:2606.08346v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving the reasoning capabilities of large language models

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures

DGX agent

arXiv:2606.08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observabili

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

CLASP: Language-Driven Robot Skill Selection and Composition using Task-Parameterized Learning

DGX agent

arXiv:2606.08169v1 Announce Type: cross Abstract: Enabling robots to understand and execute tasks from natural language commands while maintaining data efficiency remains challenging. Foundation model

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Claude Fable 5 now available on AI Gateway

DGX agent

Claude Fable 5 is now available through Vercel's AI Gateway, expanding the model options developers can access through the platform. This announcement likely details how developers can integrate and u

model-releasesvercel-blog
9 Jun 2026
Model Releases

CTS-Bench: Benchmarking Graph Coarsening Trade-offs for GNNs in Clock Tree Synthesis

DGX agent

arXiv:2602.19330v2 Announce Type: replace Abstract: Graph Neural Networks (GNNs) are increasingly explored for physical design analysis in Electronic Design Automation, particularly for modeling Clock

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Curvature-Guided LoRA: Matching Full Fine-Tuning in Function Space

DGX agent

arXiv:2603.29824v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning methods such as LoRA enable efficient adaptation of large pretrained models, but often lag behind full fine-tuning i

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Data Synthesis and Parameter-Efficient Fine-Tuning for Low-Resource NMT: A Case Study on Q'eqchi' Mayan

DGX agent

arXiv:2606.09767v1 Announce Type: cross Abstract: Neural machine translation for digitally low-resource Indigenous languages is often hindered by extreme data scarcity, prompting reliance on extractiv

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

De novo molecular generation with optical property preconditioning at the token level

DGX agent

arXiv:2606.08221v1 Announce Type: new Abstract: Designing OLED molecules with targeted optical properties remains challenging due to the scarcity of high-quality data and the limited reliability of co

model-releasesarxiv-cs-lg
9 Jun 2026
Safety

Diffuse AI Control on Fuzzy Tasks

DGX agent

arXiv:2606.08892v1 Announce Type: new Abstract: AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfiel

safetyarxiv-cs-lg
9 Jun 2026
Safety

Enhancing AI Interpretability and Safety through Localised Architectures

DGX agent

arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpre

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Evaluating Advanced Prompting on Gemini Flash for Multi-Hop Biomedical QA

DGX agent

arXiv:2606.07548v1 Announce Type: cross Abstract: The MedHopQA challenge presents a critical test for Large Language Models (LLMs): complex, multi-hop reasoning in the high-stakes biomedical domain. T

model-releasesarxiv-cs-ai
9 Jun 2026
Tutorials

Explaining Data Mixing Scaling Laws

DGX agent

arXiv:2606.08167v1 Announce Type: cross Abstract: Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical understandin

tutorialsarxiv-cs-ai
9 Jun 2026
← Previous
1…448449450451452…1357
Next →