Gabriel Willye

Data Scientist · AI Engineer · Machine Learning
📍 Campo Grande, MS, Brazil

About me

Data Scientist with 2 years of (fully remote) experience in machine learning, statistical modeling, predictive analytics and data-workflow automation. I've shipped work across government, finance and digital marketing — from an LLM fiscal-audit agent over 180M+ records to fraud-risk scoring and BI pipelines that drove real business impact. BSc in Information Systems (UFMS), with an academic track in applied AI.

PythonSQLPySparkPandasscikit-learnTensorFlowAWSGoogle CloudLooker / Power BIRAG & multi-agentNLPHDBSCAN / Random ForestDocker
Experience highlights
  • Business Analyst — Analytics · consulting (current)
  • Data Scientist — Government / tax authority · AI fiscal-audit agent (LLM+RAG) over 180M+ records, 84% accuracy
  • Data Scientist — Finance · composite asset-risk / fraud score (HDBSCAN + Random Forest), 82%
  • Business Intelligence — Digital marketing · Python BI/ETL → −25% diagnosis time, +R$2M in contracts
  • Cloud Data Engineer · end-to-end AWS pipeline (PySpark/Lambda/S3/QuickSight)
  • Academic researcher — applied AI · multi-agent systems + a published social-media study (UFMS)

Projects

Real-world and academic projects. Interactive ones are under the Live demos tab.

AI / Machine Learning

Julgador

Government · Tax-fraud detection

AI fiscal-audit agent (LLM + RAG + TensorFlow) validating invoice NCM codes across 180M+ records — 84% accuracy in the State Revenue pipeline.

Ollama Gemma 2 27B · RAG · TensorFlow

NCM Classifier

Government · Tax codes / NLP

ML/RAG classifier for Brazilian NCM tax codes — the engine feeding the Julgador audit agent.

Python · ML · RAG · TensorFlow

Asset-Risk Scoring

Finance · Fraud detection

Composite asset-risk score flagging straw-man ("laranja") fraud without banking data — HDBSCAN + Random Forest, 82% in the compliance pipeline.

HDBSCAN · Random Forest · Isolation Forest

CNN-MNIST

Deep Learning

Convolutional neural networks (LeNet / AlexNet) for handwritten-digit recognition.

TensorFlow / Keras · CNN

Abalone Age Prediction

Data Science / ML

Regression pipeline predicting abalone age; compares non-linear models (SVR best, +24.6% R² over linear).

scikit-learn · pandas

OCR + Translation

Applied AI

OCR benchmarking Tesseract vs TrOCR (CER/WER measured) + OpenCV preprocessing + auto-translation + FastAPI demo.

pytesseract · TrOCR · OpenCV · FastAPI

Face-Recognition Organizer

Applied AI · 100% offline

Organizes a photo library by who appears: phash dedup → InsightFace → clustering. ~8,886 faces → ~61 people. Code-only.

InsightFace · OpenCV · scikit-learn

Voice-ID

Applied AI · speech

Offline speaker diarization (ECAPA + clustering) + per-speaker transcription with faster-whisper.

ECAPA-TDNN · faster-whisper
Data Engineering / BI

Social Scraper

Marketing · BI / ETL

Social-data collection in a Python BI/ETL stack that powered 10 market studies & 25 diagnostics — −25% diagnosis time, +R$2M in contracts.

BeautifulSoup · Pandas · Looker Studio

Cloud Data Pipeline (AWS)

Cloud · Data Engineering · live demo

End-to-end AWS pipeline (API → PySpark/Lambda → S3 → QuickSight) + interactive Plotly dashboard (Live demos tab).

AWS · PySpark · Lambda · S3 · Plotly

Stock Data Analysis

Data Science · B3

B3 stock history via yfinance → SQLite, moving averages, Plotly viz, Prophet forecasts; real run with returns/vol/Sharpe/RSI.

yfinance · SQLite · Plotly · Prophet
Web & Applications

Movie Catalog (Filmow-style)

Web app · live demo

Search TMDB and organize films into 4 lists (seen / want-to-see / rewatch / favorites). Flask + CLI + client-side build.

Flask · JS · TMDB API

Alexandria improving

Java application

A personal media/library cataloguer in Java — being refactored into a clean version.

Java · OOP
Portfolio Demos & Analytics

Demos-Repo

15+ runnable demos

Self-contained, actually-run DS/ML/NLP/algorithm demos — churn, RFM, recommender, sentiment, A/B, market-basket, forecasting, A* and more.

Python · scikit-learn · numpy

Kaggle-Analytics

10 distinct DS problems

Ten end-to-end analyses (classification, regression, clustering, NLP, time-series, recommendation, gov EDA, credit-risk) on real public datasets. Scripts + Colab notebooks.

scikit-learn · pandas · notebooks

YouTube Watch-History Analyzer

Personal data science

Parses Takeout watch history, ranks channels, classifies ~35k videos — "Other" bucket cut 68% → 19.5%. Titles kept private.

NLP · pandas
Academic (UFMS)

UFMS Studies

BSc Information Systems

Consolidated coursework as one project repo: data structures (C/C++), databases lab (SQL+Docker), distributed computing (Flask), web programming (React), quiz app.

C/C++ · SQL · Flask · React

PET Social Networks

PET Sistemas · UFMS research

Research correlating social-media usage with psychometric scales (QSG-12) among Computing students — Pearson by segment. Aggregates only.

Python · pandas · statistics

Live demos

Interactive things you can open right here — hosted on GitHub Pages, no install.

📊 Cloud Data Dashboard

Plotly · interactive

The AWS data-pipeline analytics, re-served as a free interactive Plotly dashboard.

Plotly · Python

📈 PET Research Dashboard

Plotly · interactive

Heatmap of social-media × well-being (QSG-12) Pearson correlations, with a dropdown over 19 student segments.

Plotly · pandas

📺 YouTube History Dashboard

Plotly · interactive

35k+ watched videos broken down by category and top channels (titles kept private).

Plotly · pandas

🇧🇷 Brazil Open-Data Dashboard

Plotly · World Bank

Interactive time series of Brazil's socioeconomic indicators — GDP per capita, life expectancy, internet use, inflation, unemployment, population. Pick one from the dropdown.

Plotly · World Bank API

📑 Kaggle-Analytics Showcase

Plotly · interactive

Headline metric of each of the 10 analyses (ROC-AUC, R², ARI, accuracy…), colored by problem type — all from real runs.

Plotly

🐚 Abalone Model Comparison

Plotly · interactive

5-fold cross-validated R² of 7 regressors — SVR (RBF) best at 0.522, ~+24.6% over the linear baseline.

Plotly

📉 B3 Markets Dashboard

Plotly · Yahoo Finance

1-year cumulative return of a Brazilian (B3) stock basket plus a daily-return correlation heatmap.

Plotly · yfinance

📣 Marketing Mix (ROI)

Plotly · MMM

Adstock + regression attributing return-on-ad-spend by channel — the measure→optimize workflow of marketing-mix modeling.

Plotly · scikit-learn

🎬 Movie Catalog

TMDB · web app

Search films and sort them into seen / want-to-see / rewatch / favorites.

JS · TMDB API

🧪 Runnable demos & analyses

code · reproducible

15+ DS/ML demos and 10 dataset analyses (with Colab notebooks) — every metric produced by running the code.

Python · scikit-learn

Get in touch

Open to Data Science / AI Engineering / Data Strategy opportunities (remote-friendly).

📍 Campo Grande, MS, Brazil · 🌐 English (professional) · Portuguese (native)