Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
-
Updated
Jul 22, 2026 - Python
Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
JEPA agent playing Minecraft from pixels: latent world model + MPC planning, 664K params on one 8GB GPU, trained on raw gameplay with no labels. A complete lab notebook - including a 20-attempt research dead end, documented with its root cause.
[zenodo.20574533] CGAA: Concept Guided Adversarial Attacks
Distribution-shift teardown of openpilot v0.9.7 supercombo: does a production L2 self-driving model know when it's blind? (it doesn't, and silently)
"Publication bias and the canonization of false facts" published in eLife (2016)
Reproducible evaluation harness for hidden coordination variables in multi-agent LLM systems.
从20个项目中系统提取可复用方法论模式的实验记录。包含10轮正式审查实证(4后端)、58项发现、G5可追溯审计。不成熟框架,诚实的实验记录。
A negative result on joint-embedding predictive architectures for prediction markets. The martingale-collapse diagnostic is sound and predicts nothing, more training makes the representation worse, and the metric everything was ranked by was mostly measuring input reconstruction. Pre-registered gates, full findings record.
A $0 falsification lab across two markets — crypto (~111 hypotheses) and Polymarket prediction markets (172,830 resolved markets, 1.36M trades). 184+ techniques through one committed anti-overfitting gauntlet. 0 survive to a deployable edge; the reusable validation harness is the asset. Agent-ready. MIT.
Measurement-first research on local Mixture-of-Experts inference under a hardware contract. 3 measured laws, 4 falsified ideas, and paper site.
Research platform for discovering and rigorously falsifying crypto trading strategies: event-sourced paper trading, implementation-parity verification, and a documented negative-results record.
Synchronization Resistance: a pre-registered study measuring multi-LLM agreement difficulty. 3 pass / 1 inconclusive / 1 fail — critique wanted.
Reproducible Apple Silicon benchmark: adaptive 2/4-token prompt lookup did not beat fixed-2 on Qwen3-0.6B.
Component-level benchmark for catastrophic forgetting in world models. Two negative results: forgetting does not follow the labelled task-distance axis, and it happens in the encoder, where the usual metrics cannot see it. 375 runs, with code, data and paper.
A rigorous negative result: no pre-publish feature predicts YouTube Shorts engagement above chance, and the one 95% model is a leakage trap.
Signed attention for transformers — sinh/cosh attention giving weights in [-1,1] instead of softmax's positive-only, made FlashAttention/SDPA-compatible by channel doubling. Includes a from-scratch softmax-vs-SBA comparison and a documented negative result.
A negative-result study and a falsification protocol for LLM memory systems.
A global, open-source registry and standardized schema for null findings, non-significant outcomes, and failed trials in medical research. Built to eliminate publication bias and accelerate biomedical discovery.
Open-source market intelligence platform with a self-auditing research pipeline: pre-registered trials, placebo gates, published negative results, live paper track record vs SPY. FastAPI + Next.js + LightGBM.
Bounded negative results for subject-independent EEG affect evaluation under held-out controls.
Add a description, image, and links to the negative-results topic page so that developers can more easily learn about it.
To associate your repository with the negative-results topic, visit your repo's landing page and select "manage topics."