Self-hosted market intelligence pipeline -- collect, enrich, and surface trading signals from 10+ sources
park-intel is a self-hosted market intelligence pipeline. It collects articles from hourly and opt-in realtime source lanes (RSS, Hacker News, Reddit, GitHub, CLS, Eastmoney, and more), enriches them with keyword tagging and optional LLM-based relevance scoring, clusters related hourly articles into narrative events, publishes a daily finance newsletter, and serves everything through a REST API with a feed-first frontend.
Core sources work out of the box with zero API keys. Optional sources (Xueqiu, LLM tagging) activate when you add their credentials.
git clone https://github.com/zinan92/intel.git park-intel
cd park-intel
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # edit .env to add optional API keys
# Build frontend (one-time)
cd frontend && npm install && npm run build && cd ..
# Start server (API + frontend served together)
python main.py # open http://localhost:8001The built-in scheduler starts collecting automatically. Visit http://localhost:8001/health to see source status.
The realtime News Lane is an explicit trial. Keep it disabled while reviewing
source terms and enable it only with REALTIME_LANE_ENABLED=1; its News Items
remain readable in the feed but are excluded from the existing digest, LLM
tagger, event aggregation, and trading-signal inputs until convergence.
On an existing database, activate the seeded rows as a separate explicit step:
REALTIME_LANE_ENABLED=1 python scripts/activate_realtime_lane.py.
The realtime UI read model is available at GET /api/ui/realtime. It returns
rolling News Items, persisted AI triage (high_impact, watch, noise, or
unknown), affected assets, conditional scenarios, and source health.
SEC filings outside the configured 72-hour realtime lookback are retained as
reversible backfill records. They are excluded from realtime AI triage and the
default rolling response; use GET /api/ui/realtime?include_backfill=true for
research and retrospective triage work.
Run as background service (macOS):
bash scripts/install-service.sh # auto-starts on boot, restarts on crash
bash scripts/service-status.sh # check if running
bash scripts/uninstall-service.sh # stop and removeThe Finance Daily Newsletter is published by an external daily automation, not by the long-running collector scheduler. Keep the park-intel service focused on continuous collection, tagging, and event aggregation; schedule scripts/publish_finance_daily_newsletter.py separately at 08:00 Asia/Shanghai to generate the rolling 24-hour brief, archive the markdown output, send it to Feishu, and include a source status block so delivery issues are visible in the same push.
Configure delivery in .env:
OBSIDIAN_FINANCE_NEWSLETTER_DIR=/Users/wendy/park-io/007_finance daily newsletter
FEISHU_BOT_WEBHOOK=https://open.feishu.cn/open-apis/bot/v2/hook/...
PARK_INTEL_SKIP_FEISHU=0LLM generation uses DeepSeek first. Any DeepSeek failure retries the same frozen prompt through the local Codex CLI in an ephemeral read-only process with external tools and live search disabled. If both providers fail, no previous brief is delivered; the run is recorded as failed and the Feishu send is skipped.
Run manually without sending Feishu:
PARK_INTEL_SKIP_FEISHU=1 PYTHONPATH=. python scripts/publish_finance_daily_newsletter.py --no-generateRecover a missed historical Daily archive without changing the current brief or sending Feishu:
PYTHONPATH=. python scripts/publish_finance_daily_newsletter.py --for-date 2026-08-26Generate, archive, and send immediately:
PYTHONPATH=. python scripts/publish_finance_daily_newsletter.py| Source | Key Required | Env Var | Notes |
|---|---|---|---|
| RSS Feeds (50+) | No | -- | Blogs, newsletters, tech and crypto media |
| Hacker News | No | -- | Algolia API, score >= 20 filter |
| No | -- | 13 subreddits via RSS | |
| GitHub Trending | No | -- | Keyword-filtered trending repos |
| Yahoo Finance | No | -- | Ticker news via yfinance |
| Google News | No | -- | Query-driven news aggregation |
| GitHub Releases | Optional | GITHUB_TOKEN |
Increases API rate limit |
| Xueqiu (Chinese market) | Yes | XUEQIU_COOKIE |
Chinese market KOL commentary |
| Social KOL | Optional | -- | Requires clawfeed CLI installed |
| LLM Tagging | Optional | ANTHROPIC_API_KEY |
AI relevance scoring + narrative tags |
| CLS Telegraph | Explicit opt-in | REALTIME_LANE_ENABLED=1 |
Public rolling market-news endpoint; trial only |
| Eastmoney 7x24 | Explicit opt-in | REALTIME_LANE_ENABLED=1 |
Public fast-news endpoint; trial only |
| SEC EDGAR Watchlist | Explicit opt-in | REALTIME_LANE_ENABLED=1, SEC_EDGAR_USER_AGENT |
Official filings for the pinned 20-company watchlist and approved forms |
| BlockBeats Newsflash | Free account + explicit opt-in | BLOCKBEATS_API_KEY, REALTIME_LANE_ENABLED=1 |
Official Pro API; 5-minute free-tier baseline; secondary evidence requires confirmation |
Without ANTHROPIC_API_KEY, articles still collect and get keyword tags -- they just won't have LLM-based relevance scores or narrative tags.
Source Registry (DB)
|
v
Adapters --> Collectors (fetch + dedup) --> SQLite
|
Keyword Tagger (13 categories) |
Ticker Extractor ($NVDA, etc.) |
v
LLM Tagger (optional)
relevance_score 1-5
narrative_tags
|
v
Event Aggregator (48h window)
cross-source clustering
signal scoring
|
v
FastAPI REST API
/api/* + /api/ui/*
|
+-----------+-----------+-----------+
| | | |
React UI Quant Bridge User Health
Feed + price impact profiles Dashboard
Events from ext. topic /health
service weights
Data flow: Sources are registered in a database table (not config files). The scheduler runs one job per source type. Collectors fetch, deduplicate, and auto-tag articles on ingest. An optional LLM tagger scores relevance and generates narrative labels. The event aggregator clusters articles sharing the same narrative tag within 48-hour windows, computing a signal score (source count x avg relevance).
The /health endpoint shows per-source status including last collection time, article counts, error rates, and volume anomalies. When running the frontend, navigate to the health page to see a visual overview.
Optional: run park-intel as a persistent background service using launchd. The service auto-restarts on crash.
./scripts/install-service.sh # installs LaunchAgent and starts the service
./scripts/service-status.sh # check if the service is running
./scripts/uninstall-service.sh # stop and remove the serviceLogs go to the logs/ directory with automatic rotation.
| Endpoint | Description |
|---|---|
GET /api/health |
Per-source health status (registry-driven) |
GET /api/articles/latest |
Recent articles ?limit=20&source=rss&min_relevance=4 |
GET /api/articles/search |
Keyword search ?q=bitcoin&days=7 |
GET /api/articles/digest |
Articles grouped by source with top tags |
GET /api/articles/signals |
Topic heat + narrative momentum ?hours=24 |
GET /api/articles/sources |
Historical source statistics |
| Endpoint | Description |
|---|---|
GET /api/ui/feed |
Priority-scored feed ?user=myname&window=24h |
GET /api/ui/realtime |
Realtime rolling feed with AI triage buckets |
GET /api/ui/items/{id} |
Article detail with related items |
GET /api/ui/topics |
Topic list |
GET /api/ui/sources |
Active source list |
GET /api/ui/search |
Frontend search ?q=openai |
| Endpoint | Description |
|---|---|
GET /api/events/active |
Active events ranked by signal score |
GET /api/events/{id} |
Event detail with article timeline + price impacts |
GET /api/events/history |
Closed events archive ?tag=btc&days=30 |
| Endpoint | Description |
|---|---|
POST /api/users |
Create user profile |
GET /api/users/{username} |
Get user profile and topic weights |
PUT /api/users/{username}/weights |
Update topic weights (0.0-3.0 per topic) |
# Run tests
pytest tests/
# Run in development mode (auto-reload on file changes)
PARK_INTEL_DEV=1 python main.py
# Run collectors manually
python scripts/run_collectors.py # all sources
python scripts/run_collectors.py --source reddit # single source
# Run LLM tagger
python scripts/run_llm_tagger.py --limit 10 # score 10 unscored articles
python scripts/run_llm_tagger.py --backfill # backfill historical articles
# Backfill ticker extraction
python scripts/backfill_tickers.py
# Publish Finance Daily Newsletter
python scripts/publish_finance_daily_newsletter.py
# Run a bounded recovery for one Weekly lookback (historical Daily is archive-only)
python scripts/recover_finance_newsletters.py \
--week-ending 2026-08-30 \
--affected-start 2026-08-26Daily publication fails closed when any deduplicated candidate lacks a valid
1-5 relevance score. The failure is recorded under the Obsidian newsletter
directory's .delivery-manifests/ and does not reuse the previous published
brief. Before current publication, incomplete scoring triggers a bounded
same-window preflight; DeepSeek failures use the isolated Codex CLI fallback, with provider and
coverage recorded in the Daily archive. Weekly publication requires seven
validated Daily archives; the recovery command preserves the SQLite backup,
rescoring window, archive replacements, Weekly Feishu delivery, replay result,
and residual gaps in a JSON receipt under docs/.
GET /api/briefs/latest also exposes status, provider, scoring coverage, and
is_stale/is_usable so API consumers cannot mistake an older or unvalidated
brief for current production output. Weekly generation failures are recorded
in the Weekly delivery manifest and emit one status alert when Feishu delivery
is configured.
main.py # FastAPI entry point (port 8001)
config.py # Source seed data, collector config, env loading
scheduler.py # Registry-driven APScheduler
sources/ # Source registry, adapters, seeding
collectors/ # 10 source-type collectors (BaseCollector pattern)
events/ # Event aggregation (48h clustering, narratives)
tagging/ # Keyword tagger, LLM tagger, ticker extractor
users/ # User profiles and topic weights
bridge/ # Quant bridge (price impact from external service)
config/ # Pinned Park Exposure Registry snapshot
api/ # REST API routes
db/ # SQLAlchemy models, migrations, database init
frontend/ # React + TypeScript + Vite frontend
scripts/ # Management and utility scripts
scripts/publish_finance_daily_newsletter.py # Daily newsletter delivery
tests/ # 290+ pytest tests
- Fork the repository
- Create a feature branch (
git checkout -b feature/my-feature) - Make your changes and add tests
- Run the test suite (
pytest tests/) - Commit and push (
git push origin feature/my-feature) - Open a Pull Request
MIT
