I built this portfolio project to turn six historical NBA datasets into an auditable ETL, a canonical SQL Server model and a six-page Power BI report. The current release completes the NBA-I1 through NBA-I4 remediation plan.
The report lets you compare descriptive patterns in the available sample:
- historical win rates and scoring evolution;
- franchise age and observed performance;
- home/away scoring differences;
- shooting efficiency, turnovers and season-to-season variability;
- player physical profiles, offensive context, win streaks and recent performance.
I do not use these results to predict future games, revenue, playoff qualification or investment returns.
| Check | Result |
|---|---|
| Versioned inputs | 6 CSV · 30,638,984 LF-normalized bytes |
| Input rows | 161,111 |
| Accepted source rows | 160,956 |
| Quarantined duplicates | 155 |
| Generated historical team references | 53 |
| Canonical output rows | 161,009 |
| Unique games retained | 65,642 |
| Cross-table orphan checks | 0 |
| Core ETL test coverage | 89.92% |
The wider source folder was approximately 2.31 GB when it was audited. It is not the volume processed here, and this repository does not claim a 22 GB run. Read the source and scope note for the distinction.
six versioned CSV inputs
↓ contract v1.0.0 + typed validation
Python ETL ──→ rejects/ + manifest + SHA-256 + reconciliation
↓
canonical CSV model
↓ transactional load
SQL Server: core + audit + analytics
↓ DirectQuery
Power BI template + versionable pbi-tools source
The pipeline does not delete valid fact rows to force foreign keys. It preserves every accepted game and adds explicit historical-team records when the current 30-team dimension has no matching identifier.
You need Python 3.12. SQL Server and Power BI Desktop are only required for the BI layer.
The extracted Power BI source contains generated paths longer than the legacy Windows limit. If you use Git for Windows, enable long paths during the first checkout and keep the setting in this clone:
git -c core.longpaths=true clone https://github.com/HeKoXCode/nba-analytics-platform.git
Set-Location nba-analytics-platform
git config core.longpaths trueGitHub Desktop users can run git config core.longpaths true from Repository → Open in Command Prompt/PowerShell after cloning. If the initial checkout already failed, delete only that incomplete clone and repeat the command above in a shorter parent directory.
py -3.12 -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install -r requirements.lock
python -m pip install -e .
python -m nba_pipeline transform `
--input-dir CODE\data_raw `
--run-dir .artifacts\my-first-runEach --run-dir is immutable. Choose a new directory for every execution. A successful run contains canonical data, rejects, structured logs, the exact contract snapshot and checksummed manifests.
To run the automated suite:
python -m pip install -r requirements-dev.lock
python -m pip install -e .
python -m ruff check src tests scripts/generate_model_artifacts.py scripts/update_powerbi_project_i4.py
python -m pytest --cov=nba_pipeline --cov-report=term-missingCopy the variable names from CODE/.env.example into your own environment and provide a local secret. I do not commit credentials.
$env:NBA_SQL_DRIVER = "ODBC Driver 17 for SQL Server"
$env:NBA_SQL_SERVER = ".\SQLEXPRESS"
$env:NBA_SQL_DATABASE = "NBA_Project"
$env:NBA_SQL_TRUSTED_CONNECTION = "yes"
$env:NBA_SQL_ENCRYPT = "no"
$env:NBA_SQL_TRUST_SERVER_CERTIFICATE = "yes"
python -m nba_pipeline load-sql `
--run-dir .artifacts\my-first-run `
--evidence-dir evidence\my-first-sql-loadThe loader applies the idempotent schema, loads all canonical tables in one transaction, creates the analytical views and verifies every object and column required by Power BI. The published PBIT uses the same generic .\SQLEXPRESS source. CI overrides these local values and repeats the integration against an ephemeral SQL Server 2022 container.
Open Analisis_NBA_BestTeam.pbit after loading the local database. I organized the report into:
- Historia y evolución;
- Eficiencia y consistencia;
- Talento y perfil;
- Rachas y actualidad;
- Metodología y cierre;
- plus the cover page.
Every analytical title states its period or sample and unit. The final page records scope, quality decisions, limitations and the last integral validation date. The adjacent Analisis_NBA_BestTeam/ directory is the reviewable pbi-tools project used to compile the PBIT.
- NBA-I1–I4 verification
- Technical implementation
- Canonical data dictionary
- Model and ERD
- Versioned contract
- Real-run evidence
- CI SQL reconciliation
- Local SQL Express reconciliation
- Security review
- The committed dataset is a historical analytical sample, not an official complete NBA warehouse.
- The 53 generated team records make historical identifiers explicit; they do not invent franchise metadata.
- Duplicate rows remain inspectable in
rejects/; they are not silently discarded. - Power BI uses DirectQuery to
.\SQLEXPRESS, so you must load SQL Server Express before refreshing it. - Source access, redistribution terms and NBA-related rights must be checked before reuse.
I maintain and publish the current project as HeKoXCode. Existing Git history and notebook evidence also identify work by Percy Ignacio Marzoratti Hill and Lucas Roca. I preserve only responsibilities supported by that evidence in CONTRIBUTORS.md.
src/nba_pipeline/ production ETL, validation, watcher and SQL loader
tests/ deterministic fixture and automated checks
contracts/ generated schema contract
CODE/data_raw/ six versioned source CSV files
CODE/SQL/ database, canonical model, views and reconciliation
CODE/Dashboard.../ PBIT and extracted Power BI project
DOCS/ implementation, model, scope, security and evidence notes
evidence/ lightweight reproducibility evidence; canonical data stays ignored