中文
Data Lead · Data Engineering · GenAI in production

John Tung童曉瑜 · Hsiao-Yu Tung

Data Lead with a data-engineering spine

Ten years in data engineering, two and a half of them leading a Data Team. Most recently I shipped an enterprise Data/AI platform end-to-end for executives, marketing leads and newsroom editors — none of them engineers. Where I add the most value: getting the same result on a fraction of the budget (an inherited pipeline went from $300 to under $1 a month); putting AI into production with real quality control (when output quality slips, the system stops it rather than letting it quietly degrade); and explaining technical work in terms executives and users can act on.

Experience

Want Want China Times Media GroupData Lead · Principal Data Engineer
Apr 2026 – present
  • Built the group's cross-platform data platform from zero: Facebook, Instagram, Threads, YouTube and GA4 ingestion — 34 staggered Cloud Run Jobs writing append-only history into BigQuery, modelled with dbt. Cloud Scheduler + Cloud Run Jobs instead of Composer to control cost; dbt source freshness, account-level checks and Cloud Monitoring catch "job succeeded but wrote nothing" silent failures.
  • Rewrote an inherited revenue pipeline, $300 → under $1 a month; shipped three production tools for non-engineers: a group revenue dashboard (C-level), a social-performance dashboard (marketing and editors) and a tool that typesets monthly revenue tables into print-ready newspaper PDFs. Google sign-in and a three-tier ACL let managers add and remove users themselves.
  • Generative AI in production, with guardrails: BigQuery AI.GENERATE (Gemini) classifies comment sentiment and post topics and writes commentary grounded in current figures only; Vertex embeddings + KMEANS cluster and name negative-comment themes; a comment RAG layer is gated by a weekly LLM-as-judge eval that fails the job when quality drops. A four-tool newsroom agent is live — no free-form SQL, mandatory self-check, hard cost caps, weekly replay eval.
  • Set up how the team governs and delivers: data-governance inventory, cost-attribution labels, an access-request list negotiated with IT, decision records, staged onboarding, a cross-project knowledge base and a project briefing for executives and new engineers; a repo contract layer (no-go zones, acceptance commands, guards) so people and AI agents work under the same rules.
Cleflex TechnologyData Team Lead
May 2023 – Oct 2025
  • Led a four-person Data Team — scheduling, prioritisation and delivery quality — running Scrum (stand-ups, sprint planning, retros) on Jira/GitLab, and partnering with product to turn data into Superset dashboards used for decisions.
  • Airflow crawling ETL (Scrapy → S3/MongoDB), scheduled Elasticsearch reads feeding AI reviews, Telegram alerts; built an Apache Iceberg + Spark lakehouse for distributed analysis and standardised deployment with Docker/makefiles so teammates could stand up environments fast.
Amber TechnologyData Engineer
Nov 2022 – Apr 2023
  • Automated on-prem → GCP Cloud Storage ETL with Fluentd, daily Airflow computations, introduced PySpark for distributed processing and planned the overall data architecture.
PIXNETSenior Algorithm Engineer
Sep 2021 – Nov 2022
  • Fluentd log ETL on GKE → Cloud Storage/BigQuery; Search Console API collection at 9M rows/day; trending-topic, merchant-ranking and recommendation algorithms; SEO (Sitemap, JSON-LD).
LiTVSoftware Engineer
Oct 2019 – Sep 2021
  • Ansible-deployed Fluentd fleets (GCE + on-prem IDC) streaming into BigQuery; AWS Data Pipeline cross-cloud ETL; automated Data Studio reporting; gender prediction for anonymous users and customer algorithms (XGBoost/RandomForest/SVM).
EarlierNCHC Software Engineer · NTUST Programming Lecturer · Dataa Data Analyst
Oct 2015 – Aug 2019
  • TWCC AI-platform acceptance testing and CKAN data-market automation; taught Python data analysis; delivered public-sector and enterprise sentiment-analysis projects (crawl → MongoDB → text mining → client reports).

Skills

Data engineering
BigQuery · dbt · Cloud Run Jobs + Scheduler · Airflow · PySpark/Spark · Apache Iceberg · Fluentd · Scrapy/Selenium
Cloud
GCP (Cloud Run · Secret Manager · Monitoring · GKE · GCE · GCS) · AWS (Data Pipeline · S3) · Azure · on-prem ↔ cloud
AI / GenAI
Gemini via BigQuery AI.GENERATE · Vertex embeddings + KMEANS · RAG + LLM-as-judge · agent guardrails & cost control · XGBoost/LightGBM · Chinese NLP
Delivery
Streamlit (OIDC + ACL) · Superset · Mann-Kendall/Sen's slope · Docker · Cloud Build/GitHub Actions · Ansible · Scrum · Jira/GitLab · Python · SQL

Education

NTUST — Information Management
PhD candidate (candidacy 2019; left for industry) 2016 – 2019
Yuan Ze Univ. — Information Management
M.S. · Thesis on noise filtering & word-of-mouth mining in car forums 2013 – 2015
Kainan Univ. — Information & E-Commerce
B.S. 2009 – 2013
Updated 2026-09-10 · https://myps6415.github.io/cv