Research
Bullynx research and evaluation methods
Reproducible methods for evaluating AI chart readers: what to test, how to score it, what counts as a refusal, and which limitations must remain visible.
Protocol published · Results not yet published
AI chart-reader evaluation methodology
A fixed 100-point rubric for market structure, level fidelity, evidence, refusal behavior, and repeated-run consistency, with an explicit protocol-only publication status.
Read the methodology →What the protocol scores
An AI chart reader is judged on what a second person can verify from the same screenshot, not on whether a call made money. Five dimensions carry twenty points each, scored 0, 10 or 20 against fixed written anchors so two evaluators reach the same number for the same output.
- Market structure. Does it name the visible trend or range and every material swing that was labeled in advance, without inventing a swing that is not on the chart?
- Level fidelity. Are the support, resistance and invalidation prices within a tolerance fixed before the test: 0.5% of the displayed price range, or one labeled axis increment, whichever is larger? The invalidation must also sit on the correct side of price.
- Evidence and uncertainty. Is every material conclusion tied to something visible in the image, and is each pre-labeled ambiguity disclosed rather than smoothed over? Citing an indicator that is not in the screenshot scores zero.
- Refusal behavior. On a cropped price axis or a blurred image, does it decline to give exact numbers, say which information is missing, and ask for a usable replacement? Over-refusing a chart that was readable costs points as well, so caution is not a free score.
- Repeated-run consistency. Sent the identical screenshot again, does it keep the same regime, the same invalidation direction and levels within tolerance? A tool that changes its mind between two identical inputs is not measuring the chart.
The six required inputs are fixed too: a readable trend, a readable range, a chart crowded with overlapping indicators, a chart with the price axis cropped off, a blurred image, and one screenshot submitted repeatedly. Ground truth is written down before any product output is read, so the answer key cannot move to fit a result.
Publication rule
Bullynx will separate protocols from results. A methodology can be published before a benchmark exists; a score cannot. Any future result must include the chart set, tested version, date, raw denominator, repeated runs, limitations, and conflict-of-interest disclosure.