The Data Analyst's Toolkit — First Session (Sports Business)

The kickoff session for a student building a portfolio of end-to-end sports-business-operations analytics projects for internship interviews. It introduces the five-tool stack - SQL, Python, Power BI, Excel, and GitHub - giving the honest pros and cons of each, showing where each one fits in a project, and explaining how they hand off to one another to form a single pipeline that extracts, analyzes, presents, and publishes. Across 31 slides there are five tool sections, each with a strengths-and-limits breakdown and a tool-choice trap, two runnable examples (a ticket-revenue SQL query and an attendance EDA in pandas), an SVG pipeline diagram, the end-to-end recipe, a full worked sports project carried through all five tools, and two checks. No prior knowledge is assumed.

Subject: Business Analytics · 59 slides · applied lesson

Open the interactive version of this deck · Homework for this lesson

What this lesson covers

The lesson, slide by slide

1. The Data Analyst's Toolkit

Title

First Session · Sports Business Analytics

SQL, Python, Power BI, Excel, and GitHub — what each one is for, where each falls short, and how they snap together into portfolio projects that get you the internship.

2. What you'll walk away with today

Objectives

This is a map session, not a deep dive. We're laying out the whole toolkit so every future session has a place to hang. By the end you can:

3. What survived from Decision Analysis — Oceanview Property Purchase?

Warm-up

Discussion prompt

Before we open The Data Analyst's Toolkit — First Session (Sports Business): without looking back, what was the main idea of Decision Analysis — Oceanview Property Purchase, and what could you do by the end of it that you could not do before?

Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.

Answer:

Master's-level walkthrough of the Oceanview Development property-purchase case (Camm et al., 4th ed., p. Builds the decision tree, folds it back to the no-research decision (bid; EV = $50,000), revises the referendum odds with Bayes' theorem from the survey's reliability, derives the conditional bid/don't-bid strategy (EV $229,268 if approval predicted, −$74,576 if rejection), and values the $15,000 study with EVSI = $44,000 against EVPI = $70,000 (efficiency 62.9%).

4. The Big Picture

Section

Part 1

5. What we're actually building

Concept

The goal isn't "learn five tools." It's 2–3 end-to-end projects you can show in an interview. An interviewer doesn't want a chart — they want to watch you take a messy business question and turn it into a decision.

end-to-end project — One project that spans the whole arc: a real question → the data pulled to answer it → the analysis → a presentation a non-technical boss understands → published proof anyone can review.

  1. A real operations question ("do promo nights grow paid attendance?")
  2. The data, pulled and shaped to answer exactly that
  3. The analysis — what's actually true in the numbers
  4. A dashboard that tells the story to a decision-maker
  5. A public write-up: question, method, finding

6. One question, five tools

Concept

Each tool owns one stage. They're not competitors — they're a relay team. Here's the whole map on one slide:

ToolStageIts one job
SQLExtractPull the exact rows you need from the database
PythonAnalyzeClean, explore, and forecast
Power BIPresentInteractive dashboards for stakeholders
ExcelPresent / modelFast dashboards & models everyone can open
GitHubPublishHost the code + write-up interviewers click

7. Fill in: Stage for One question, five tools

Comparison

Comparison matrix

From One question, five tools: refill the Stage column from what you know. The rest of the table is as it appeared.

ToolStageIts one job
SQLExtractPull the exact rows you need from the database
PythonAnalyzeClean, explore, and forecast
Power BIPresentInteractive dashboards for stakeholders
ExcelPresent / modelFast dashboards & models everyone can open
GitHubPublishHost the code + write-up interviewers click

8. A relay race, not five separate races

Intuition

Picture a 4×100 relay. No single runner wins it. Each runs one leg, then hands off the baton — and the baton here is your data.

SQL runs the first leg and hands a clean table to Python. Python runs its leg and hands a result to Power BI or Excel. They hand the finished story to GitHub, where the world sees the race.

Drop the baton at any handoff — a broken export, a messy column, a project that only lives on your laptop — and the whole run doesn't count. The handoffs are the skill.

9. Break it if you can: A relay race, not five separate races

Counterexample

Discussion prompt

Picture a 4×100 relay. No single runner wins it. Each runs one leg, then hands off the baton — and the baton here is your data.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

Answer:

SQL runs the first leg and hands a clean table to Python. Python runs its leg and hands a result to Power BI or Excel. They hand the finished story to GitHub, where the world sees the race.

10. SQL — Get the Data

Section

Part 2 · Extract

11. What SQL is

Concept

SQL is how you talk to a database. You describe the rows you want — filtered, joined, and summed — and it hands them back. It is the front door to almost every real dataset in a sports org: ticketing, CRM, concessions, sponsorship.

query — A single statement that says which columns, from which tables, filtered how, grouped how, ordered how. One query is a repeatable, shareable definition of "the data."

12. Take the definitions apart: end-to-end project vs query

Definition probe

Sort into buckets

Every line below is part of the definition of end-to-end project or of query — one or the other, never both. Put each where it belongs.

end-to-end project
One project that spans the whole arc; a real question → the data pulled to answer it → the analysis → a presentation a non-technical boss understands → published proof anyone can review.
query
A single statement that says which columns, from which tables, filtered how, grouped how, ordered how.; One query is a repeatable, shareable definition of "the data."
b1
One project that spans the whole arc: a real question → the data pulled to answer it → the analysis → a presentation a non-technical boss understands → published proof anyone can review.
b2
A single statement that says which columns, from which tables, filtered how, grouped how, ordered how. One query is a repeatable, shareable definition of "the data."

13. SQL — strengths & limits

Concept

✅ Great at:

⚠️ Where it stops:

14. What has to happen first: SQL in action: revenue per game by channel

Ranking

Put in order

Put the moves of SQL in action: revenue per game by channel into the order they have to happen.

  1. FROM ticket_sales WHERE season = 2024
  2. GROUP BY channel
  3. SELECT the aggregates, then ORDER BY rev_per_game DESC

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Point at the right table and keep only last season's rows — before doing any math.

15. SQL in action: revenue per game by channel

Worked example

Question: last season, which sales channel drove the most ticket revenue per game?

SELECT
    channel,
    COUNT(DISTINCT game_id)                       AS games,
    SUM(tickets_sold)                             AS tickets,
    ROUND(SUM(revenue) / COUNT(DISTINCT game_id)) AS rev_per_game
FROM ticket_sales
WHERE season = 2024
GROUP BY channel
ORDER BY rev_per_game DESC;

FROM ticket_sales WHERE season = 2024

Why: Point at the right table and keep only last season's rows — before doing any math. Filtering first means the database scans far less data.

GROUP BY channel

Why: Collapse thousands of individual sales into one row per channel (Box Office, Web, Resale, Group Sales), so each aggregate is computed within a channel.

SELECT the aggregates, then ORDER BY rev_per_game DESC

Why: Sum revenue and divide by distinct games to get a fair per-game figure, then sort so the top channel is the first row you read.

channelgamesrev_per_game
Group Sales41$182,400
Box Office41$164,900
Web41$151,250
Resale41$38,600

16. Fill in: rev_per_game for SQL in action: revenue per game by channel

Comparison

Comparison matrix

From SQL in action: revenue per game by channel: refill the rev_per_game column from what you know. The rest of the table is as it appeared.

channelgamesrev_per_game
Group Sales41$182,400
Box Office41$164,900
Web41$151,250
Resale41$38,600

17. Something is wrong here: doing the crunching in the wrong place

Anomaly

Predict first

A student writes this, and it looks reasonable:

Pull everything and figure it out later.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Excel caps near a million rows and chokes long before that; you've moved the database's own job onto a tool that can't do it.

Let SQL do what it's built for; export only the answer.

Why: Excel caps near a million rows and chokes long before that; you've moved the database's own job onto a tool that can't do it.

18. Trap: doing the crunching in the wrong place

Trap

The trap

Pull everything and figure it out later.

SELECT * FROM ticket_sales;
-- 9.4 million rows -> Excel/Python
-- then filter, group, and total by hand

Drag 9.4M raw rows out, then group them in Excel

Why: Excel caps near a million rows and chokes long before that; you've moved the database's own job onto a tool that can't do it.

The fix

Let SQL do what it's built for; export only the answer.

SELECT channel, ROUND(SUM(revenue)) AS rev
FROM ticket_sales
WHERE season = 2024
GROUP BY channel;
-- 4 rows out

Aggregate at the source, pull 4 tidy rows

Why: Move the computation to the data. What leaves the database is small, clean, and ready for the next tool in the relay.

19. Inspect it line by line: Trap: doing the crunching in the wrong place

Error analysis

Annotate

Walk the callouts on Trap: doing the crunching in the wrong place. Each one is a place this is easy to get subtly wrong.

  • Excel caps near a million rows and chokes long before that; you've moved the database's own job onto a tool that can't do it.
  • Move the computation to the data. What leaves the database is small, clean, and ready for the next tool in the relay.

20. Python — Explore & Forecast

Section

Part 3 · Analyze

21. What Python does here

Concept

Python (with pandas) is your workshop for the messy middle: clean the SQL pull, slice it every which way to find what's true, and — when the question calls for it — build a simple forecast.

EDA (exploratory data analysis) — Poking the data to see what's there before you commit to an answer: distributions, group averages, outliers, trends. This is where most real insight is found.

22. Python — strengths & limits

Concept

✅ Great at:

⚠️ Where it stops:

23. What has to happen first: Python in action: attendance by day of week

Ranking

Put in order

Put the moves of Python in action: attendance by day of week into the order they have to happen.

  1. read_csv(...) the tidy pull from SQL
  2. groupby("weekday")["attendance"].mean()
  3. sort_values(ascending=False) and print

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The relay handoff: Python starts from the small, clean table SQL exported — not from the raw database.

24. Python in action: attendance by day of week

Worked example

Question: which game days pull the biggest crowds — so operations can staff and price accordingly?

import pandas as pd

games = pd.read_csv("games_2024.csv")

by_day = (games
          .groupby("weekday")["attendance"]
          .mean()
          .sort_values(ascending=False))

print(by_day.round(0))

read_csv(...) the tidy pull from SQL

Why: The relay handoff: Python starts from the small, clean table SQL exported — not from the raw database.

groupby("weekday")["attendance"].mean()

Why: Split the season by day of week, then average attendance within each — the split-apply-combine move that is the heart of pandas EDA.

sort_values(ascending=False) and print

Why: Rank the days so the pattern reads instantly: the answer, not just the numbers.

weekdayavg attendance
Saturday18,240
Friday16,980
Sunday15,110
Wednesday11,430

25. Fill in: avg attendance for Python in action: attendance by day of week

Comparison

Comparison matrix

From Python in action: attendance by day of week: refill the avg attendance column from what you know. The rest of the table is as it appeared.

weekdayavg attendance
Saturday18,240
Friday16,980
Sunday15,110
Wednesday11,430

26. Something is wrong here: using Python for the whole project

Anomaly

Predict first

A student writes this, and it looks reasonable:

Do everything in Python, including the final report.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: You spend the afternoon fighting fonts and legends to rebuild what Power BI does in three clicks — and the VP still can't filter it themselves.

Use each tool for its leg of the relay.

Why: You spend the afternoon fighting fonts and legends to rebuild what Power BI does in three clicks — and the VP still can't filter it themselves.

27. Trap: using Python for the whole project

Trap

The trap

Do everything in Python, including the final report.

Write 60 lines of matplotlib to hand-style the executive bar chart

Why: You spend the afternoon fighting fonts and legends to rebuild what Power BI does in three clicks — and the VP still can't filter it themselves.

The fix

Use each tool for its leg of the relay.

Explore & forecast in Python; export the clean result to Power BI/Excel

Why: Python finds the truth; the presentation tool makes it interactive and boardroom-ready. Right tool, right stage.

28. Power BI — Show the Story

Section

Part 4 · Present

29. What Power BI is

Concept

Power BI turns a finished dataset into an interactive dashboard: charts, KPI cards, and slicers a non-technical stakeholder clicks to explore themselves. It's the deliverable a front-office exec actually remembers.

✅ Great at: live filtering (by team, month, channel), auto-refreshing from a source, and a professional look with almost no design work.

⚠️ Where it stops: it's a display layer, not an analysis engine — it shows what you already found. Full features are Windows-first, and it won't fix bad inputs.

30. The dashboard is the handshake

Intuition

Your SQL and Python are the engine under the hood. Nobody in the front office ever sees the engine. They see the dashboard — that's your whole analysis, wearing a suit.

So the dashboard has one job: let a busy decision-maker reach your finding in ten seconds, then poke it ("what about weekend games? just the rivals?") and trust what comes back.

31. Break it if you can: The dashboard is the handshake

Counterexample

Discussion prompt

Your SQL and Python are the engine under the hood. Nobody in the front office ever sees the engine. They see the dashboard — that's your whole analysis, wearing a suit.

That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.

Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.

32. Something is wrong here: a pretty dashboard is not proof

Anomaly

Predict first

A student writes this, and it looks reasonable:

Plug raw data straight into Power BI and ship the good-looking result.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: Duplicated rows and a mislabeled channel are invisible once they're a smooth line.

Verify upstream; let the dashboard show a story you already trust.

Why: Duplicated rows and a mislabeled channel are invisible once they're a smooth line. Garbage in, beautiful garbage out — and you present it with confidence.

33. Trap: a pretty dashboard is not proof

Trap

The trap

Plug raw data straight into Power BI and ship the good-looking result.

Skip the checking because the charts look clean

Why: Duplicated rows and a mislabeled channel are invisible once they're a smooth line. Garbage in, beautiful garbage out — and you present it with confidence.

The fix

Verify upstream; let the dashboard show a story you already trust.

Clean & sanity-check in SQL/Python first, then visualize

Why: The dashboard is the last mile, not the analysis. It should only ever display numbers you've already confirmed are real.

34. Break it on purpose: a pretty dashboard is not proof

Break the constraint

Discussion prompt

The rule this trap just fixed:

The dashboard is the last mile, not the analysis. It should only ever display numbers you've already confirmed are real.

Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?

Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.

Answer:

Duplicated rows and a mislabeled channel are invisible once they're a smooth line. Garbage in, beautiful garbage out — and you present it with confidence.

35. Excel — The Universal Tool

Section

Part 5 · Present / Model

36. What Excel is for (still)

Concept

Excel is the language every business person already speaks. For quick dashboards, a scenario model, or a table you can hand to anyone, it's the fastest tool on Earth — no install, no login.

✅ Great at: instant pivots and charts, what-if models, and universal shareability — everyone can open an .xlsx.

⚠️ Where it stops: it breaks down past ~1M rows, and hand-built workbooks are error-prone and hard to reproduce or audit.

37. Something is wrong here: making Excel do it by hand every week

Anomaly

Predict first

A student writes this, and it looks reasonable:

Rebuild the report by copy-pasting fresh data in each week.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: No record of what you did, one wrong paste corrupts it silently, and next month you can't explain — or reproduce — how the number was made.

Keep the transformation in SQL/Python; let Excel just present the output.

Why: No record of what you did, one wrong paste corrupts it silently, and next month you can't explain — or reproduce — how the number was made.

38. Trap: making Excel do it by hand every week

Trap

The trap

Rebuild the report by copy-pasting fresh data in each week.

Paste new numbers over last week's, re-drag the formulas

Why: No record of what you did, one wrong paste corrupts it silently, and next month you can't explain — or reproduce — how the number was made.

The fix

Keep the transformation in SQL/Python; let Excel just present the output.

Excel consumes a clean, regenerated export

Why: The repeatable steps live in code; Excel is the friendly last-mile view. Re-run the query, refresh the sheet — same result every time.

39. GitHub — Publish the Proof

Section

Part 6 · Publish

40. What GitHub is

Concept

GitHub hosts your code and write-up at a public link, with a full history of every change. For your goal it matters for one blunt reason: it's the portfolio interviewers actually click.

repository (repo) — A project's home on GitHub: its code, its data-handling scripts, and — most importantly — its README, the front page that explains what the project is and what you found.

✅ Great at: proving the work is yours, showing your process, and version history. ⚠️ Where it stops: it's not analysis, and a repo with no README is a locked door.

41. Term to definition: The Data Analyst's Toolkit — First Session (Sports Business)

Matching

Match the pairs

Match each term to the definition this lesson gave it — not the one you would guess from the word.

  • t1. query
  • t2. EDA (exploratory data analysis)
  • t3. repository (repo)
  • d1. A single statement that says which columns, from which tables, filtered how, grouped how, ordered how. One query is a repeatable, shareable definition of "the data."
  • d2. Poking the data to see what's there before you commit to an answer: distributions, group averages, outliers, trends. This is where most real insight is found.
  • d3. A project's home on GitHub: its code, its data-handling scripts, and — most importantly — its README, the front page that explains what the project is and what you found.

Why: These are the working definitions of query, EDA (exploratory data analysis), repository (repo) as The Data Analyst's Toolkit — First Session (Sports Business) uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.

42. Something is wrong here: a repo nobody can read (or should see)

Anomaly

Predict first

A student writes this, and it looks reasonable:

Push everything — data, secrets, no explanation.

It is wrong. Say what breaks — and say it before you turn the page.

Correct: You've leaked a credential, blown the size limit, and left a reviewer staring at files with no idea what the project even is.

Publish the story, not the mess.

Why: You've leaked a credential, blown the size limit, and left a reviewer staring at files with no idea what the project even is. That's a strike, not a portfolio piece.

43. Trap: a repo nobody can read (or should see)

Trap

The trap

Push everything — data, secrets, no explanation.

Commit a 500 MB data dump, a DB password, and 12 loose files

Why: You've leaked a credential, blown the size limit, and left a reviewer staring at files with no idea what the project even is. That's a strike, not a portfolio piece.

The fix

Publish the story, not the mess.

Add a .gitignore for data & secrets; write a README (question, method, finding)

Why: A stranger lands on the page and gets it in 30 seconds. That readable repo is the link you paste into every application.

44. Which of these survive contact with The Data Analyst's Toolkit — First Session…?

Two truths and a lie

Sort into buckets

Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.

Holds up
Each tool owns one stage. They're not competitors — they're a relay team. Here's the whole map on one slide:; Picture a 4×100 relay. No single runner wins it. Each runs one leg, then hands off the baton — and the baton here is your data.; Your SQL and Python are the engine under the hood. Nobody in the front office ever sees the engine. They see the dashboard — that's your whole analysis, wearing a suit.
Breaks
Pull everything and figure it out later.; Do everything in Python, including the final report.
sound
These are stated as this lesson states them — each one survives the edge cases The Data Analyst's Toolkit — First Session (Sports Business) puts it through.
flawed
Each of these is lifted from a trap in this deck: reasonable-sounding, and wrong in a way that only shows up once you rely on it.

45. Putting It Together

Section

Part 7

46. Picture it first: The five handoffs

Picture it

Figure (svg): A left-to-right pipeline: a Business question feeds SQL, which feeds Python, which feeds Power BI and Excel, which feeds GitHub.

Extract → analyze → present → publish. One direction, five handoffs.

Discussion prompt

Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.

Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.

Answer:

Here's the whole relay on one track. The data is the baton; each tool runs one leg and passes it on.

47. The five handoffs

Concept

Here's the whole relay on one track. The data is the baton; each tool runs one leg and passes it on.

Figure (svg): A left-to-right pipeline: a Business question feeds SQL, which feeds Python, which feeds Power BI and Excel, which feeds GitHub.

Extract → analyze → present → publish. One direction, five handoffs.

48. Rebuild the recipe: The end-to-end recipe

Ranking

Put in order

These are the steps of The end-to-end recipe, scrambled. Put them back in order before the next slide shows you.

  1. Ask a real operations question — one sentence, decision-oriented.
  2. SQL: pull and pre-aggregate just the rows that answer it.
  3. Python: clean, explore, and (if useful) forecast.
  4. Power BI / Excel: build the dashboard that tells the story to a non-technical stakeholder.
  5. GitHub: publish the code + a README so anyone can follow it.
  6. Write it up: the question, what you found, what you'd do about it.

Why: This is the order the recipe itself gives. Recalling the sequence without the slide in front of you is the difference between recognising the method and being able to run it — most of what goes wrong in practice is a step done out of turn.

49. The end-to-end recipe

Pattern

Every project you build this summer follows the same six steps. Memorize the shape — the topic changes, the shape doesn't:

  1. Ask a real operations question — one sentence, decision-oriented.
  2. SQL: pull and pre-aggregate just the rows that answer it.
  3. Python: clean, explore, and (if useful) forecast.
  4. Power BI / Excel: build the dashboard that tells the story to a non-technical stakeholder.
  5. GitHub: publish the code + a README so anyone can follow it.
  6. Write it up: the question, what you found, what you'd do about it.

50. Where does it stop working: The end-to-end recipe

Edge cases

Discussion prompt

The end-to-end recipe works on the cases you have just seen. Push it to the edge: what is the most degenerate input it still handles — empty, zero, one item, everything equal — and what is the first case where it stops being true? Name the case, not just "it breaks".

Hint: Try the smallest legal input, then the largest, then the one where two things collide. Methods are specified at their edges; the middle takes care of itself.

Answer:

Every project you build this summer follows the same six steps. Memorize the shape — the topic changes, the shape doesn't:

51. What has to happen first: One project, all five tools

Ranking

Put in order

Put the moves of One project, all five tools into the order they have to happen.

  1. SQL — pull per-game paid attendance, promo flag, opponent, and price tier for three seasons
  2. Python — compare promo vs. non-promo games, adjusting for opponent strength and weekend effects
  3. Power BI / Excel — dashboard: attendance lift per promo type, filterable by opponent and month
  4. GitHub — publish the query, the notebook, and a README stating question, method, and finding

Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The warehouse holds millions of ticket rows; SQL filters to the ~120 home games and returns one tidy row per game — the baton.

52. One project, all five tools

Worked example

The question the front office actually asks: do our promo nights (giveaways, theme nights) grow paid attendance — or just move fans from one game to another?

SQL — pull per-game paid attendance, promo flag, opponent, and price tier for three seasons

Why: The warehouse holds millions of ticket rows; SQL filters to the ~120 home games and returns one tidy row per game — the baton.

Python — compare promo vs. non-promo games, adjusting for opponent strength and weekend effects

Why: A raw average is misleading (promos cluster on weekends vs. weak opponents). pandas controls for that, and can forecast next season's lift.

Power BI / Excel — dashboard: attendance lift per promo type, filterable by opponent and month

Why: The VP won't read your notebook — they'll click a dashboard. This is the deliverable they remember and share.

GitHub — publish the query, the notebook, and a README stating question, method, and finding

Why: This is the link you paste into an application. It turns "I know analytics" into "here's the proof."

Five tools, one baton, one story. That is an end-to-end project — and you'll have two or three of them.

53. Work backwards from the answer: One project, all five tools

Reverse engineer

Discussion prompt

Work backwards. The example finished here:

GitHub — publish the query, the notebook, and a README stating question, method, and finding

What was it asked to do, and what must it have been given? Reconstruct the problem from its answer.

Hint: Every quantity in the result had to enter somewhere. Account for each one.

Answer:

The question the front office actually asks: do our promo nights (giveaways, theme nights) grow paid attendance — or just move fans from one game to another?

54. Rule out three: Check yourself: pick the first tool

Elimination

Eliminate the wrong options

You need last season's ticket sales for the three home rivals, already summed per game, out of a 10-million-row database. Which tool is the right first stop?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. SQL — filter and aggregate at the source, export the small result
  • B. Excel — open the database and pivot the 10M rows
  • C. Power BI — connect and build the chart first
  • D. GitHub — create the repo first

Survives elimination: A

Why: Extraction is the first stage, and SQL owns it. Filtering to three opponents and summing per game happens on the server, so only a handful of clean rows ever leave the database — exactly what the next tool wants.

55. Check yourself: pick the first tool

Check

Think about the stage before you answer.

Check your understanding

You need last season's ticket sales for the three home rivals, already summed per game, out of a 10-million-row database. Which tool is the right first stop?

  • A. SQL — filter and aggregate at the source, export the small result (correct)
  • B. Excel — open the database and pivot the 10M rows
  • C. Power BI — connect and build the chart first
  • D. GitHub — create the repo first

Answer: A

Why: Extraction is the first stage, and SQL owns it. Filtering to three opponents and summing per game happens on the server, so only a handful of clean rows ever leave the database — exactly what the next tool wants.

Why B tempts people
Excel can't hold 10M rows (it caps near ~1.05M) and shouldn't be the one filtering a database — that's SQL's job. Excel comes later, on the small result.
Why C tempts people
Power BI is the present stage. Pointing it at 10M raw rows with no upstream aggregation is slow and skips the analysis you haven't done yet.
Why D tempts people
GitHub is the publish stage — it hosts finished work. There's nothing to publish before you've pulled and analyzed the data.

56. Rule out three: Check yourself: make it reviewable

Elimination

Eliminate the wrong options

Your cleaning steps and forecast live in a Jupyter notebook on your laptop. Which move best turns that into portfolio proof an interviewer can actually evaluate?

3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.

  • A. Push the notebook plus a README (question, method, finding) to a public GitHub repo
  • B. Email them a screenshot of the final dashboard
  • C. Keep it local and offer to demo it live sometime
  • D. Paste the Python code into the Power BI report's description box

Survives elimination: A

Why: Publishing is the last stage, and GitHub owns it. A public repo with a README lets a stranger see the question, follow your method, and trust the finding on their own time — that's what "reviewable" means and it's the link you send.

57. Check yourself: make it reviewable

Check

An interviewer says, "show me your work." What makes the project reviewable?

Check your understanding

Your cleaning steps and forecast live in a Jupyter notebook on your laptop. Which move best turns that into portfolio proof an interviewer can actually evaluate?

  • A. Push the notebook plus a README (question, method, finding) to a public GitHub repo (correct)
  • B. Email them a screenshot of the final dashboard
  • C. Keep it local and offer to demo it live sometime
  • D. Paste the Python code into the Power BI report's description box

Answer: A

Why: Publishing is the last stage, and GitHub owns it. A public repo with a README lets a stranger see the question, follow your method, and trust the finding on their own time — that's what "reviewable" means and it's the link you send.

Why B tempts people
A screenshot shows the result but hides the work — the interviewer can't check the method, the cleaning, or whether the number is even right.
Why C tempts people
Work that only lives on your laptop isn't a portfolio; a live-only demo can't be reviewed before the interview and disappears after it.
Why D tempts people
Power BI's description box isn't a code host — no history, no README, no real way to read or run a notebook. That's what GitHub is for.

58. Connect it up: The Data Analyst's Toolkit — First Session (Sports Business)

Connect it up

Draw it

One page, no notation unless you need it: draw how these connect — The Big Picture · SQL — Get the Data · Python — Explore & Forecast · Power BI — Show the Story · Excel — The Universal Tool · GitHub — Publish the Proof. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.

59. Recap — the toolkit, in one breath

Recap

You can now place every tool in the relay and say what it's for:

StageToolThe one thing it does
ExtractSQLPull the exact rows from the database
AnalyzePythonClean, explore, forecast
PresentPower BI / ExcelDashboards a stakeholder understands
PublishGitHubThe code + write-up interviewers click

Next session: we pick your first project's question and write the SQL that pulls its data — leg one of the relay.

Sources

  1. Mode — SQL Tutorial (SELECT, WHERE, GROUP BY, aggregate functions)
  2. pandas — Group by: split-apply-combine
  3. Microsoft — What is Power BI?
  4. GitHub Docs — About READMEs / Ignoring files
  5. SQL query and pandas snippet both executed against sample sports-ticketing data before shipping. — Author verification run, 2026-07-14 (first-session intro deck).

Want this taught 1-on-1? Alexander tutors Business Analytics — $55/session, free consultation.

Book on Wyzant · Text (657) 465-8108