Files

13 KiB
Raw Permalink Blame History

Invitation to Independent Replication

Version: 1.0 Date: 2025-12-14 Purpose: Formal invitation to third-party verification of experimental claims


I. STATEMENT OF INTENT

We actively invite independent replication of our results.

This is not a courtesy—it's a challenge. We provide:

  • Complete source code
  • Full experimental data
  • Step-by-step replication protocols
  • Exact hardware/software specifications
  • Expected results with falsification thresholds

If you cannot reproduce our claimed 0.00% CV in algorithmic variance, we want to know immediately.


II. WHY REPLICATE?

For Skeptics

You think our claims are too good to be true?

Prove it. Replicate our experiment and report differing results. We'll either:

  1. Acknowledge a bug/error in our methodology, OR
  2. Help you identify environmental differences

Either outcome advances science.


For Supporters

You want to use this in your own work?

Validate it first. Independent replication:

  • Strengthens the evidence
  • Identifies edge cases
  • Builds community trust

For Researchers

You're working on adaptive systems?

Replication provides:

  • Baseline for comparison
  • Validation of statistical methods
  • Understanding of failure modes

III. REPLICATION LEVELS

We offer three tiers of replication, from easy to comprehensive:

Level 1: Quick Smoke Test (15 minutes)

Goal: Verify basic functionality and 0% CV claim

Protocol:

# Clone repository
git clone https://github.com/rajames440/StarForth.git
cd StarForth

# Build fastest target
make fastest

# Run single DoE trial
./build/amd64/fastest/starforth --doe --config=C_FULL

# Expected output:
# Cache hit rate: 17.39 ± 0.00%
# Runtime: ~7-10 ms ± 60% (environment-dependent)

Success Criteria:

  • Build completes without errors
  • 780+ tests pass
  • DoE run produces cache CV ≈ 0%

Failure Threshold: Cache CV > 0.5%


Level 2: Partial Replication (4 hours)

Goal: Reproduce convergence claim (25.4% improvement)

Protocol:

# Build
make fastest

# Run 30 trials for C_FULL configuration
for i in {1..30}; do
  ./build/amd64/fastest/starforth --doe --config=C_FULL > results/run_${i}.csv
done

# Analyze convergence
python3 scripts/analyze_convergence.py results/

# Expected output:
# Early runs (1-15): ~10.20 ms
# Late runs (16-30): ~7.61 ms
# Improvement: ~25.4%
# p-value: < 0.001

Success Criteria:

  • All 30 runs show cache CV = 0.00%
  • Late runs show statistically significant improvement (p < 0.05)
  • Improvement magnitude within [20%, 30%]

Failure Threshold:

  • Cache CV > 0.1% in any run
  • No convergence (p > 0.05)
  • Improvement < 10%

Level 3: Full Replication (2-3 days)

Goal: Reproduce all 5 formal claims (C1-C5)

Protocol:

# Build all configurations
make fastest

# Run 90 trials (3 configs × 30 runs)
./scripts/full_replication.sh

# Statistical validation
Rscript scripts/statistical_validation.R

# Expected outputs:
# - C1: Algorithmic CV = 0.00%
# - C2: Convergence = 25.4 ± 1.2%
# - C3: Variance separation (0% vs 70%)
# - C4: Reproducibility (all runs identical)
# - C5: p < 10⁻³⁰ for determinism

Success Criteria:

  • All 5 claims reproduce within error margins
  • Checksums match expected values
  • Statistical tests yield same significance levels

Failure Threshold: Any claim fails to reproduce (see FORMAL_CLAIM_TABLE.md)


IV. EXACT REPLICATION (Bit-for-Bit)

Purpose: Eliminate environmental differences

Protocol:

# Build Docker image
docker build -t starforth-replication -f Dockerfile.replication .

# Run full experiment
docker run --rm -v $(pwd)/results:/results starforth-replication

# Verify checksums
cd results/
sha256sum -c EXPECTED_CHECKSUMS.txt

# Expected: All checksums match

Success Criteria:

  • All output files match SHA256 checksums
  • Bit-for-bit identical results to original

Why This Matters: If Docker replication fails, the issue is in our code, not environment.


Manual Exact Replication

Requirements:

  • OS: Ubuntu 22.04.3 LTS (kernel 6.2.0-39-generic)
  • GCC: 11.4.0
  • Make: 4.3
  • CPU: x86_64 with AVX2 support
  • RAM: 16GB minimum
  • Disk: 10GB free space

Setup:

# Disable frequency scaling
sudo cpupower frequency-set -g performance

# Disable Turbo Boost
echo 1 | sudo tee /sys/devices/system/cpu/intel_pstate/no_turbo

# Disable ASLR
echo 0 | sudo tee /proc/sys/kernel/randomize_va_space

# Set process affinity
taskset -c 0 ./starforth --doe

Expected Variability:

  • Cache CV: 0.00% (exact match)
  • Runtime: ±20% due to hardware differences (acceptable)
  • Convergence rate: ±5% (acceptable)

V. REPORTING RESULTS

Successful Replication

If you reproduce our results, please report:

Template:

## Replication Report

**Replicator**: [Your Name / Institution]
**Date**: [YYYY-MM-DD]
**Replication Level**: [1/2/3]
**Commit SHA**: [git commit hash]

### Results
- Cache CV: [your value] (expected: 0.00%)
- Convergence: [your value] (expected: 25.4%)
- Statistical significance: p = [your p-value] (expected: p < 0.001)

### Hardware
- CPU: [model]
- RAM: [size]
- OS: [version]

### Conclusion
✅ Successfully replicated all claims within error margins.

**Contact**: [your email]

Where to Report: File GitHub issue with tag replication-success


Failed Replication

If you CANNOT reproduce our results, please report:

Template:

## Replication Failure Report

**Replicator**: [Your Name / Institution]
**Date**: [YYYY-MM-DD]
**Replication Level**: [1/2/3]
**Commit SHA**: [git commit hash]

### Observed Deviations
- **Claim**: [which claim failed, e.g., C1: Determinism]
- **Expected**: [our claimed value]
- **Observed**: [your value]
- **Deviation**: [percentage/absolute difference]

### Environment
- CPU: [model]
- RAM: [size]
- OS: [version]
- GCC: [version]
- Build flags: [from Makefile]

### Logs
[Attach: build logs, runtime logs, error messages]

### Attempted Mitigations
[What you tried to fix it]

**Contact**: [your email]

Where to Report: File GitHub issue with tag replication-failure

Our Commitment: We will respond within 48 hours with either:

  1. Acknowledgment of bug - If we can reproduce your failure
  2. Diagnostic questions - If we suspect environmental differences
  3. Request for data - To analyze deviations

VI. COMMON REPLICATION ISSUES

Issue 1: "I got 0.05% CV instead of 0.00%"

Diagnosis: Likely measurement precision (acceptable)

Resolution:

  • Our claim: CV < 0.1% (see FORMAL_CLAIM_TABLE.md)
  • Your result: 0.05% CV is within threshold
  • Conclusion: Replication successful

Issue 2: "My convergence is 18%, not 25%"

Diagnosis: Likely hardware differences (acceptable)

Resolution:

  • Our claim: Convergence exists (p < 0.05)
  • Your result: 18% is statistically significant
  • Conclusion: Replication successful (magnitude varies with hardware)

Why: Convergence rate depends on relative cost of dictionary lookup (CPU-specific)


Issue 3: "I got 70% CV in cache decisions"

Diagnosis: CRITICAL - Replication failed

Resolution:

  1. Verify build: make clean && make fastest
  2. Check for random number generators in code (should be none)
  3. Verify ASLR disabled: cat /proc/sys/kernel/randomize_va_space (should be 0)
  4. Run under valgrind: valgrind --tool=memcheck ./starforth --doe
  5. Report to us immediately (GitHub issue)

This is a genuine failure—we want to know.


Issue 4: "Build fails with linker errors"

Diagnosis: Likely GCC version mismatch

Resolution:

# Check GCC version
gcc --version  # Should be 11.x

# If different, use Docker (see Section IV)
docker build -t starforth .
docker run starforth

VII. BOUNTY FOR FALSIFICATION

We offer recognition to the first replicator who:

  1. Falsifies Claim C1 (Determinism)

    • Reproduces CV > 0.1% under controlled conditions
    • Co-authorship on erratum paper
  2. Falsifies Claim C2 (Convergence)

    • Shows no convergence (p > 0.05) across 30 runs
    • Acknowledgment in future publications
  3. Identifies Critical Bug

    • Bug that invalidates core claims
    • Named credit in bug fix commit

Why: Science advances through falsification. If our claims are wrong, we want to know.


VIII. CROSS-INSTITUTIONAL REPLICATION

We Seek Collaborators

Ideal Replication Partners:

  • Academic labs - Publish independent validation
  • Industry teams - Validate for production use
  • Skeptical researchers - Best critics make best validators

What We Provide:

  • Technical support (email/video chat)
  • Access to original hardware (if needed)
  • Co-authorship on replication study (if desired)

Contact: rajames440@gmail.com (Robert A. James)


IX. REPLICATION TIMELINE

Phase 1: Initial Replications (3 months)

Goal: 3-5 independent replications at Level 2+

Target: Academic institutions, open-source contributors

Deliverable: Replication reports published as GitHub issues


Phase 2: Cross-Platform Validation (6 months)

Goal: Validate on ARM, RISC-V, non-Linux platforms

Target: Embedded systems, L4Re deployments

Deliverable: Platform-specific replication guides


Phase 3: Long-Term Monitoring (ongoing)

Goal: Track replication success rate over time

Metric: % of attempts that successfully reproduce claims

Target: > 90% success rate (indicates robust methodology)


X. FAQ

Q: "Do I need permission to replicate?"

A: No. Code is CC0 (public domain). Replicate freely.


Q: "Can I use this in my own research?"

A: Yes. If replication succeeds, cite our work. If it fails, publish your findings.


Q: "What if I find a bug?"

A: File a GitHub issue. We'll fix it and credit you.


Q: "What if my hardware is different?"

A: Acceptable. Report your results with hardware specs. Convergence magnitude may vary; determinism (0% CV) should not.


Q: "Can I modify the code?"

A: Yes (CC0 license). But for replication, use unmodified code first.


Q: "What if I can't reproduce AND can't identify why?"

A: Contact us (rajames440@gmail.com). We'll help diagnose.


XI. REPLICATION SCORECARD

We will maintain a public scorecard of replication attempts:

Replicator Institution Date Level Result Notes
R.A. James Original 2025-12-08 3 Pass Baseline
[Your Name] [Your Org] [Date] [1/2/3] [/] [Link to report]

Transparency: All replication attempts (success or failure) will be documented publicly.


XII. PSYCHOLOGICAL WARFARE ASPECT

Why This Document Exists

Traditional approach: "Trust us, we did the experiment."

Our approach: "Don't trust us—replicate it yourself."

Effect:

  1. Skeptics who won't replicate look weak
  2. Skeptics who try and fail expose themselves
  3. Skeptics who succeed validate our work
  4. Everyone sees we're confident enough to invite scrutiny

Outcome: Criticism becomes costly (requires effort), validation becomes compelling (independent verification).


XIII. CONCLUSION

We don't just tolerate replication—we demand it.

Three possible outcomes:

  1. You replicate successfully → Our claims are validated
  2. You find a bug → We fix it and science advances
  3. 🤷 You don't attempt replication → Your criticism carries less weight

All three outcomes are acceptable to us.

The ball is in your court. Replicate or acknowledge you haven't.


XIV. REPLICATION CHECKLIST

Before you start:

  • Read EXECUTIVE_SUMMARY.md (understand what you're replicating)
  • Read FORMAL_CLAIM_TABLE.md (know the exact claims)
  • Read NEGATIVE_RESULTS.md (understand failure modes)
  • Choose replication level (1/2/3)
  • Set up environment (Docker recommended)

During replication:

  • Document all deviations from protocol
  • Save all logs and outputs
  • Record hardware/software specs
  • Note any unusual behavior

After replication:

  • Compare results to expected values
  • Calculate deviations
  • File report (success or failure)
  • Contact us if needed

Bottom Line: If you're confident in your skepticism, replicate our experiment and prove us wrong. We'll thank you for it.

License: See ./LICENSE