Skip to content

Engineering an Automated Analysis Pipeline

learnfrc.com
learnfrc.comAuthor
Veer Bajaj
Veer BajajMaintainer

The advanced techniques only pay off if they run reliably at event pace. This lesson lays out the whole pipeline: scouting data goes in, public API data gets joined to it, validated metrics come out, and the whole thing rebuilds with one command between match cycles.

Keep the stages decoupled so that any one of them can be debugged or replaced.

  1. Ingest. Scanned QRScout rows append to a raw source, either a Google Sheet or a CSV. Append rows only, and never edit one in place.
  2. Enrich. A script pulls TBA (event_matches, event_oprs, event_rankings) and Statbotics (get_team_event) for the event, keyed by team number.
  3. Compute. Join the scouting averages with OPR and EPA, then compute your custom metrics (per-level coral rates, endgame reliability, defense rating) along with validation flags.
  4. Output. A rank table and per-match prep sheets, regenerated on demand.
import pandas as pd, tbapy, statbotics
EVENT = '2025cc'
tba = tbapy.TBA('YOUR_TBA_KEY')
sb = statbotics.Statbotics()
# 1. Ingest scouting (exported from your QRScout sheet)
raw = pd.read_csv('scouting_raw.csv')
agg = raw.groupby('teamNumber').mean(numeric_only=True)
# 2/3. Enrich with public analytics
oprs = tba.event_oprs(EVENT)['oprs'] # {'frc254': 78.3, ...}
agg['opr'] = [oprs.get(f'frc{t}') for t in agg.index]
agg['epa'] = [ sb.get_team_event(int(t), EVENT).get('epa', {}).get('unitless')
for t in agg.index ] # EPA at this event
# 4. Validation flags
agg['low_sample'] = raw.groupby('teamNumber').size() < 3
agg['scout_vs_opr_gap'] = (agg['ptsContributed'] - agg['opr']).abs()
agg.sort_values('ptsContributed', ascending=False).to_csv('rank.csv')

Use the Statbotics fields parameter and the TBA simple/keys options to keep responses small and fast. Sending the TBA ETag back as If-None-Match also means a repeated refresh returns a cheap 304 Not Modified.

  • One command, and idempotent. Running the pipeline twice gives the same result, so you can rebuild anytime without manual cleanup.
  • Fail loud on validation. If low_sample or scout_vs_opr_gap is high for a top team, surface it so you can send a super-scout, instead of silently ranking on thin data.
  • Cache API calls. Store the TBA and Statbotics responses locally on each refresh, so a flaky venue connection or a rate limit does not block your rebuild. QR scouting already works offline, so your analysis should degrade gracefully too.
  • Keep a manual fallback. If the script breaks mid-event, the agg spreadsheet alone (from the Worked Examples module) still produces a usable ranking. Never let the fancy pipeline be a single point of failure.

Put together with the earlier modules, this is a complete and defensible scouting operation: offline-first collection, measured scout accuracy, validated custom metrics, cross-checks against public analytics, predictive what-ifs, explicit defense evaluation, and a one-command rebuild that feeds picklists and pre-match briefs.

  • Decouple the pipeline into ingest (append-only) -> enrich (TBA + Statbotics) -> compute (custom metrics + validation) -> output (rank + briefs).
  • Make rebuilds idempotent and one-command, cache API calls, and fail loud on low-sample or scouting-vs-OPR gaps to trigger super-scouts.
  • Always keep a manual spreadsheet fallback so the automated pipeline is never a single point of failure at an event.

This lesson was adapted from learnfrc.com.