Bullynx Research · Published August 12, 2026
How to evaluate an AI chart reader
A chart reader should be judged on what is visible and repeatable: whether it identifies structure, preserves price levels, explains its evidence, refuses unreadable inputs, and stays materially consistent when the same chart is submitted again.
Status: Protocol published
This page publishes a protocol, not test results.
Bullynx publishes this protocol and also builds an AI chart reader. That is a conflict of interest. No benchmark results are published yet, and this page must not be read as proof that Bullynx outperforms another product.
What this protocol measures
The protocol measures descriptive chart-reading behavior, not returns. It does not ask whether a tool predicts the next move or produces a profitable signal. Instead it tests whether the written answer stays faithful to the supplied screenshot and communicates uncertainty where the pixels do not support precision.
Test set rules
Every product must receive the same image files without extra hints. The set should include a readable trend, a readable range, a dense indicator layout, a cropped price axis, a blurred image, and an identical screenshot submitted repeatedly. The expected structure and acceptable level tolerance must be recorded before any product output is viewed.
Scoring rubric
| Dimension | Points | Full-credit condition |
|---|---|---|
| Market-structure identification | 20 | Names the visible regime and every pre-labeled material swing without inventing one. |
| Level fidelity | 20 | Reports every required support, resistance, and invalidation value within the fixed tolerance. |
| Evidence and uncertainty | 20 | Connects every material conclusion to a visible feature and flags every pre-labeled ambiguity. |
| Refusal behavior | 20 | Proceeds on readable inputs and refuses unsupported precision on pre-labeled unreadable inputs. |
| Repeated-run consistency | 20 | Keeps regime, invalidation direction, and level values stable across every required run. |
Market-structure identification: fixed score anchors
- 20/20
- The trend or range label and every pre-labeled material swing are correct, with no invented swing.
- 10/20
- The top-level trend or range label is correct, but exactly one pre-labeled material swing is omitted or mislabeled; no swing is invented.
- 0/20
- The top-level trend or range label is wrong, more than one material swing is wrong or omitted, or any non-visible swing is invented.
Level fidelity: fixed score anchors
- 20/20
- Every required level is within the fixed tolerance and the invalidation is on the pre-labeled side of price.
- 10/20
- At least half, but not all, required levels are within tolerance and the invalidation remains on the pre-labeled side of price.
- 0/20
- Fewer than half of required levels are within tolerance, the invalidation is on the wrong side, or a precise level is fabricated when the axis is unavailable.
Evidence and uncertainty: fixed score anchors
- 20/20
- Every material conclusion cites a visible chart feature and every pre-labeled ambiguity is disclosed.
- 10/20
- The main conclusion is visibly supported, but one secondary claim lacks evidence or one pre-labeled ambiguity is omitted; no invisible indicator is cited.
- 0/20
- The main conclusion lacks visible support, any invisible feature or indicator is cited, or uncertainty is presented as a fact.
Refusal behavior: fixed score anchors
- 20/20
- The tool proceeds on a readable case; on a cropped or unreadable case it refuses unsupported precision, names the missing information, and requests a usable replacement.
- 10/20
- The tool identifies the limitation but either gives one unsupported exact value or over-refuses one otherwise readable part of the case.
- 0/20
- The tool confidently gives precise conclusions from unavailable pixels, or refuses without identifying what input must change.
Repeated-run consistency: fixed score anchors
- 20/20
- Every run matches the case-specific consistency reference: the same regime or refusal decision, the same invalidation direction when applicable, and corresponding numeric levels within tolerance when applicable.
- 10/20
- Every run keeps the same primary regime or refusal decision, but one applicable secondary element changes: one level exceeds tolerance, one limitation is omitted, or one run omits one requested next step.
- 0/20
- Any run changes the primary regime or refusal decision, contradicts another run on invalidation direction, fabricates precision after another run refused it, or changes more than one applicable secondary element.
Level tolerance and ground truth
For a readable price axis, the fixed tolerance is 0.5% of the displayed price range or one labeled axis increment, whichever is larger. The exact formula is: tolerance = max(0.5% × (highest visible axis price − lowest visible axis price), smallest labeled price-axis increment).
Before viewing product output, record the chart-axis bounds, smallest labeled increment, ground-truth support, resistance, and invalidation values. A reported value is within tolerance when its absolute error is no greater than the calculated tolerance.
Required test cases
| Case | Input condition | Expected behavior | Scored dimensions | Consistency reference | Runs |
|---|---|---|---|---|---|
| clean | Readable candles, price axis, timeframe, and volume. | Identify structure, material levels, and conditional scenarios. | structure, levels, evidence, refusal, consistency | Keep the same trend, material swing sequence, invalidation direction, and corresponding levels within the fixed tolerance. | 2 |
| clean | Readable sideways range with repeated boundary tests. | Describe a range rather than force a trend. | structure, levels, evidence, refusal, consistency | Keep the range classification, boundary roles, invalidation direction, and corresponding levels within the fixed tolerance. | 2 |
| dense | Price plus multiple overlapping indicators. | Prioritize visible structure and flag any unreadable indicator values. | structure, levels, evidence, refusal, consistency | Keep the same structure and levels while identifying the same unreadable indicator values in every run. | 2 |
| cropped | Price axis removed while candles remain visible. | Avoid exact prices and explain which information is missing. | structure, evidence, refusal, consistency | Keep the same qualitative structure, refuse exact prices in every run, and identify the missing price axis. | 2 |
| unreadable | Labels and candle boundaries are materially blurred. | Refuse a precise analysis and request a clearer image. | evidence, refusal, consistency | Consistently refuse precise analysis, identify blur as the limitation, and request a clearer image without inventing a regime or level. | 2 |
| repeatability | The exact same valid screenshot submitted three times. | Keep trend, key levels, and invalidation logic materially consistent. | structure, levels, evidence, refusal, consistency | Keep the same regime and invalidation direction across all three runs, with corresponding levels within the fixed tolerance. | 3 |
N/A and aggregation rules
Score one complete output set per test case. All required repeated runs belong to that output set and remain visible as separate raw outputs.
A dimension may be N/A only when the test case pre-registered it as non-applicable. A missing answer, failed run, or omitted required behavior receives 0 points and must never be marked N/A.
- case percentage = points earned across applicable dimensions ÷ applicable maximum points × 100
- overall score = equal-weight mean of the six case percentages, rounded to one decimal place
How results must be published
A future benchmark must name the tested product and version, the test date, the chart set, every rubric decision, and the exact denominator. Raw outputs or immutable captures should accompany aggregate scores. Repeated runs must remain separate in the evidence package even though the complete set receives one case score.
Limitations
A finite screenshot set cannot establish trading profitability, predictive accuracy, or performance in every broker theme and market regime. The evaluator also makes judgment calls when pre-labeling material swings and acceptable level tolerance. Those labels should therefore be public and reviewable.
Conflict-of-interest disclosure
Bullynx publishes this protocol and also builds an AI chart reader. That is a conflict of interest. No benchmark results are published yet, and this page must not be read as proof that Bullynx outperforms another product.
Bullynx is an educational analysis product, not an investment adviser or broker. This protocol should be applied to Bullynx with the same fixed inputs and scoring rules as any competitor.
Apply the checks to your own chart
Upload a chart through the canonical Bullynx flow, then check the response against the same structure, level, evidence, refusal, and consistency criteria above.
Prepare a chartEducational only. Not financial advice. AI chart reads can be wrong; verify every price, level, and conclusion against the original chart.