A result is only as reliable as the evidence beneath it.
Historical research becomes useful when its evidence survives scrutiny. This case study describes two connected but distinct workstreams: BacktestApp, a local historical simulation application, and CSL alpha research, a separate feature, model-state and portfolio-evaluation workflow. The common contribution is an engineering approach to making research inspectable while protecting proprietary decision logic.
We separate input identity, decision availability, modeled execution, financial reconciliation and research history. Archived acceptance records support early BacktestApp checkpoints. Selected source contracts and test definitions support a more detailed explanation of later modules; retained research reports inform the alpha case study. These evidence classes remain distinct.
Two workstreams.
One evidence discipline.
Historical simulation engineering
Data handling, causal rule evaluation, modeled fills, a Decimal ledger, trade evidence and research lineage in a local software workflow.
Applied research evaluation
Input certification, causal model reconstruction, saved-state scoring and capital-normalized portfolio research under separate numerical contracts.
The research process distinguishes three kinds of work: what the implementation does, what testing exercised, and what a market-facing conclusion can actually support. A source file does not by itself establish execution; a passing backtest does not by itself establish an investable edge.
Make every layer answer
one clear question.
A software architecture is useful when it makes disagreement diagnosable. Did the wrong data enter? Did a rule use observations too early? Did an order fill under unrealistic assumptions? Did fees enter accounting twice? Different questions require different evidence.
Causal first.
More precise second.
BacktestApp separates completed strategy observations from modeled execution observations. Prefix-invariance tests mutate later data and compare prior decisions. The minute-resolution execution contract uses 1-minute observations without allowing their future ranges to inform an earlier decision.
Missing observations remain visible as gaps rather than fabricated fills. When protective stop and target thresholds are reachable within the same minute, the configured ambiguity policy is recorded. This is more defensible than pretending OHLCV data expose a complete tick-by-tick exchange path.
Show that the money
adds up.
A deliberately synthetic round trip demonstrates why a single ledger matters. These prices are invented accounting inputs, not an alpha signal or a performance statistic.
1,000.00000 − 499.50000 − 0.49950 + 549.45000 − 0.54945 = 1,048.90105For the separate CSL portfolio workstream, a different numerical contract reconciles price PnL, signed funding where applicable, costs, and capital allocation. A percentage return on a trade is not automatically a contribution to portfolio equity.
When evidence disagrees,
investigate the contract.
The alpha research workstream raises the difficulty: incomplete inputs, model-state provenance, selection bias and portfolio concentration. Our approach is not to hide inconvenient findings; it is to turn them into targeted engineering questions.
Diagnose residual gaps, accept only verified recovery candidates, and retain exclusions where input evidence is insufficient.
When an earlier workflow exhibits leakage, require an availability-aware reconstruction rather than promoting its result.
Keep unavailable estimator state as a provenance blocker; any authorized replacement must carry its own identity.
Retain failed candidates and distinguish exploratory comparisons from genuinely independent future evaluation.
Evidence, classified
by what it actually proves.
The BacktestApp source and archived acceptance records document multiple development checkpoints. The figures below are the reported passing-test counts for separate archived checkpoints, not a single current test run and not cumulative.
| Evidence class | What it supports |
|---|---|
| Archived acceptance records | 36, 56, 59 tests recorded across distinct milestones; local smoke workflow and deterministic output hashes. |
| Inspected source contracts | Stage 3 execution, Stage 4 evidence, and Stage 5 lineage/governance behavior as described in source and defined tests. |
| Saved alpha reports | Input-audit and matched-comparison summaries; these are not independent raw-source reruns. |
| Bounded toolkit record | Sixteen passing checks and five saved-ledger reconciliations are described in the revision's author-retained verification register; independently checkable output is not published here. |
| Public synthetic fixture | Decimal arithmetic identities shown above, independently reproducible without strategy disclosure. |
Project evidence for private code and alpha remains under controlled review. Source presence, historic acceptance, fresh checks and independent certification are not interchangeable claims.
Research questions become
reviewable deliverables.
We build toward a decision, not a chart. The output should let a reviewer understand assumptions, reproduce supported observations and see unresolved items before committing money or architecture.
Price, funding and cost units made explicit; cashflows matched or discrepancies recorded.
Timestamp/availability conventions and focused future-mutation checks.
Capital, cost, coverage and eligibility assumptions aligned between candidate and control.
Versioned report, scoped reproduction instructions and engineering findings.
Defined scope.
Useful engineering.
Available now: an offline Evidence Toolkit Starter designed to review five saved research experiments, check package identity, export reports and examine CSV cashflow/timestamp consistency. It is not an arbitrary-strategy backtesting engine or an order-execution product.
Commissioned engineering: integration with accepted client datasets, broader temporal or execution stress, customized ledger checks and decision-ready technical reports. Each engagement defines the question, data access, acceptance criteria and deliverables before execution.
Publication boundary: this technical edition shares architecture responsibilities, synthetic arithmetic and general lessons. Feature combinations, private model state, thresholds, exact signal times and trade sequences remain outside the public material. The BacktestApp spot simulation and separate CSL numerical portfolio replay use distinct accounting contracts.
Methodological context.
Relevant literature informs why selection bias and backtest overfitting matter. The references below contextualize the research discipline; the paper does not claim to have estimated PBO or deflated Sharpe for its proprietary workstreams.
- Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2014). Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance. Notices of the AMS, 61(5), 458–471. Read source ↗
- Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2017). The Probability of Backtest Overfitting. Journal of Computational Finance, 20(4), 39–69. Read source ↗
- Bailey, D. H., & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting, and Non-Normality. Journal of Portfolio Management, 40(5), 94–107. Open DOI ↗
Self-authored technical case study prepared from archived BacktestApp acceptance records, selected source/test inspection and retained CSL summaries. It is not represented as independent peer review. The public edition is designed to explain engineering value without publishing reconstructible proprietary trading logic.
Want your research
to withstand better questions?
Bring a problem, a dataset or a result worth examining. Start with a focused engineering review.
Discuss a quant engineering project ↗