About me
Data Scientist with 2 years of (fully remote) experience in machine learning, statistical modeling, predictive analytics and data-workflow automation. I've shipped work across government, finance and digital marketing — from an LLM fiscal-audit agent over 180M+ records to fraud-risk scoring and BI pipelines that drove real business impact. BSc in Information Systems (UFMS), with an academic track in applied AI.
- Business Analyst — Analytics · consulting (current)
- Data Scientist — Government / tax authority · AI fiscal-audit agent (LLM+RAG) over 180M+ records, 84% accuracy
- Data Scientist — Finance · composite asset-risk / fraud score (HDBSCAN + Random Forest), 82%
- Business Intelligence — Digital marketing · Python BI/ETL → −25% diagnosis time, +R$2M in contracts
- Cloud Data Engineer · end-to-end AWS pipeline (PySpark/Lambda/S3/QuickSight)
- Academic researcher — applied AI · multi-agent systems + a published social-media study (UFMS)
Projects
Real-world and academic projects. Interactive ones are under the Live demos tab.
Julgador
AI fiscal-audit agent (LLM + RAG + TensorFlow) validating invoice NCM codes across 180M+ records — 84% accuracy in the State Revenue pipeline.
NCM Classifier
ML/RAG classifier for Brazilian NCM tax codes — the engine feeding the Julgador audit agent.
Asset-Risk Scoring
Composite asset-risk score flagging straw-man ("laranja") fraud without banking data — HDBSCAN + Random Forest, 82% in the compliance pipeline.
CNN-MNIST
Convolutional neural networks (LeNet / AlexNet) for handwritten-digit recognition.
Abalone Age Prediction
Regression pipeline predicting abalone age; compares non-linear models (SVR best, +24.6% R² over linear).
OCR + Translation
OCR benchmarking Tesseract vs TrOCR (CER/WER measured) + OpenCV preprocessing + auto-translation + FastAPI demo.
Face-Recognition Organizer
Organizes a photo library by who appears: phash dedup → InsightFace → clustering. ~8,886 faces → ~61 people. Code-only.
Voice-ID
Offline speaker diarization (ECAPA + clustering) + per-speaker transcription with faster-whisper.
Social Scraper
Social-data collection in a Python BI/ETL stack that powered 10 market studies & 25 diagnostics — −25% diagnosis time, +R$2M in contracts.
Cloud Data Pipeline (AWS)
End-to-end AWS pipeline (API → PySpark/Lambda → S3 → QuickSight) + interactive Plotly dashboard (Live demos tab).
Stock Data Analysis
B3 stock history via yfinance → SQLite, moving averages, Plotly viz, Prophet forecasts; real run with returns/vol/Sharpe/RSI.
Movie Catalog (Filmow-style)
Search TMDB and organize films into 4 lists (seen / want-to-see / rewatch / favorites). Flask + CLI + client-side build.
Alexandria improving
A personal media/library cataloguer in Java — being refactored into a clean version.
Demos-Repo
Self-contained, actually-run DS/ML/NLP/algorithm demos — churn, RFM, recommender, sentiment, A/B, market-basket, forecasting, A* and more.
Kaggle-Analytics
Ten end-to-end analyses (classification, regression, clustering, NLP, time-series, recommendation, gov EDA, credit-risk) on real public datasets. Scripts + Colab notebooks.
YouTube Watch-History Analyzer
Parses Takeout watch history, ranks channels, classifies ~35k videos — "Other" bucket cut 68% → 19.5%. Titles kept private.
UFMS Studies
Consolidated coursework as one project repo: data structures (C/C++), databases lab (SQL+Docker), distributed computing (Flask), web programming (React), quiz app.
PET Social Networks
Research correlating social-media usage with psychometric scales (QSG-12) among Computing students — Pearson by segment. Aggregates only.
Live demos
Interactive things you can open right here — hosted on GitHub Pages, no install.
📊 Cloud Data Dashboard
The AWS data-pipeline analytics, re-served as a free interactive Plotly dashboard.
📈 PET Research Dashboard
Heatmap of social-media × well-being (QSG-12) Pearson correlations, with a dropdown over 19 student segments.
📺 YouTube History Dashboard
35k+ watched videos broken down by category and top channels (titles kept private).
🇧🇷 Brazil Open-Data Dashboard
Interactive time series of Brazil's socioeconomic indicators — GDP per capita, life expectancy, internet use, inflation, unemployment, population. Pick one from the dropdown.
📑 Kaggle-Analytics Showcase
Headline metric of each of the 10 analyses (ROC-AUC, R², ARI, accuracy…), colored by problem type — all from real runs.
🐚 Abalone Model Comparison
5-fold cross-validated R² of 7 regressors — SVR (RBF) best at 0.522, ~+24.6% over the linear baseline.
📉 B3 Markets Dashboard
1-year cumulative return of a Brazilian (B3) stock basket plus a daily-return correlation heatmap.
📣 Marketing Mix (ROI)
Adstock + regression attributing return-on-ad-spend by channel — the measure→optimize workflow of marketing-mix modeling.
🎬 Movie Catalog
Search films and sort them into seen / want-to-see / rewatch / favorites.
🧪 Runnable demos & analyses
15+ DS/ML demos and 10 dataset analyses (with Colab notebooks) — every metric produced by running the code.
Get in touch
Open to Data Science / AI Engineering / Data Strategy opportunities (remote-friendly).
📍 Campo Grande, MS, Brazil · 🌐 English (professional) · Portuguese (native)