Research

Bullynx research and evaluation methods

Reproducible methods for evaluating AI chart readers: what to test, how to score it, what counts as a refusal, and which limitations must remain visible.

Protocol published · Results not yet published

AI chart-reader evaluation methodology

A fixed 100-point rubric for market structure, level fidelity, evidence, refusal behavior, and repeated-run consistency, with an explicit protocol-only publication status.

Read the methodology →

What the protocol scores

An AI chart reader is judged on what a second person can verify from the same screenshot, not on whether a call made money. Five dimensions carry twenty points each, scored 0, 10 or 20 against fixed written anchors so two evaluators reach the same number for the same output.

  • Market structure. Does it name the visible trend or range and every material swing that was labeled in advance, without inventing a swing that is not on the chart?
  • Level fidelity. Are the support, resistance and invalidation prices within a tolerance fixed before the test: 0.5% of the displayed price range, or one labeled axis increment, whichever is larger? The invalidation must also sit on the correct side of price.
  • Evidence and uncertainty. Is every material conclusion tied to something visible in the image, and is each pre-labeled ambiguity disclosed rather than smoothed over? Citing an indicator that is not in the screenshot scores zero.
  • Refusal behavior. On a cropped price axis or a blurred image, does it decline to give exact numbers, say which information is missing, and ask for a usable replacement? Over-refusing a chart that was readable costs points as well, so caution is not a free score.
  • Repeated-run consistency. Sent the identical screenshot again, does it keep the same regime, the same invalidation direction and levels within tolerance? A tool that changes its mind between two identical inputs is not measuring the chart.

The six required inputs are fixed too: a readable trend, a readable range, a chart crowded with overlapping indicators, a chart with the price axis cropped off, a blurred image, and one screenshot submitted repeatedly. Ground truth is written down before any product output is read, so the answer key cannot move to fit a result.

Publication rule

Bullynx will separate protocols from results. A methodology can be published before a benchmark exists; a score cannot. Any future result must include the chart set, tested version, date, raw denominator, repeated runs, limitations, and conflict-of-interest disclosure.