About Me
Welcome to my website! I'm Calazans Macchiutti, a Data Scientist with a Ph.D. in Physics (CBPF) and a strong background in statistical modeling, scientific reasoning, and applied machine learning.
Today, I work as a Data Scientist at ESPM, where I build data solutions that support strategic academic decision-making, including predictive analytics, KPI development, automation pipelines, and dashboards. I enjoy turning complex and messy data into clear insights, reliable metrics, and scalable products that improve real-world processes.
My academic training gave me a rigorous foundation in problem-solving, experimental design, and quantitative analysis. Over time, I transitioned that expertise into industry-focused work, applying machine learning to time series, anomaly detection, and performance forecasting — always aiming for practical impact, clarity, and measurable results.
Areas of Focus
- Machine Learning & Predictive Modeling: classification, forecasting, anomaly detection
- Time Series Analytics: pattern recognition, monitoring, efficiency indicators
- Applied Data Science: dashboards, KPI frameworks, data quality metrics
- Automation & Data Pipelines: scalable workflows for analytics and reporting
- AI for Decision Support: insights that improve processes and operations
Professional Experience
- Data Scientist — ESPM (2024–Present)
- Senior Data Scientist — BINAHKI (2023–2024)
- Ph.D. Researcher — CBPF (2020–2024)
I'm especially motivated by projects where data can drive better decisions, optimize resources, and improve outcomes. If you're looking for a Data Scientist who combines technical depth, strong analytical thinking, and a practical product-driven mindset, feel free to reach out.
Academic Profile
Research in Condensed Matter Physics and Materials Science
Research Focus
During my doctoral program, I focused my studies on the physical properties of transition metal oxides, with a special emphasis on the investigation of double perovskite-type structures and their impact on the "exchange bias" phenomenon.
My research covers a variety of techniques:
- Magnetic measurements
- Electrical and thermal transport studies
- X-ray diffraction
- Magnetoresistance analysis
- Investigation of strong electronic correlation phenomena
Materials Synthesis Experience
Extensive experience in materials synthesis through various methods:
- Sol-gel synthesis
- Co-precipitation methods
- Solid-state synthesis
- Floating zone techniques
All syntheses included both polycrystals and single crystals of ternary and quaternary rare earth oxides, as well as binary single crystals such as V₂O₅.
Characterization Techniques
- X-ray photoelectron spectroscopy (XPS)
- Raman spectroscopy
- Atomic force microscopy (AFM)
Early Research
- Musical Acoustics and Environmental Acoustics
- Applications of electromagnetic waves in magnetic resonance
- Analysis of sound waves in music
View all publications →
Data Science
Professional Summary
My journey has been shaped by the development and application of statistical models, with a strong focus on identifying patterns, understanding trends, and generating actionable insights.
I developed a strong foundation in scientific problem-solving, quantitative modeling, and rigorous analytical reasoning through end-to-end research projects. This experience includes hypothesis formulation, experimental planning, data acquisition, and the interpretation of complex multivariate phenomena. With this background, I present below a detailed overview of my professional trajectory from a technical perspective.
Technical Skills
Programming & Data Manipulation
- Python (Pandas, Polars, NumPy)
- SQL for querying, transformation, and data management
- Dask and PySpark for large-scale processing
Data Visualization & Statistical Analysis
- Matplotlib, Seaborn, Plotly
- SciPy and Statsmodels for statistical modeling and inference
Machine Learning & AI
- TensorFlow, Keras, Scikit-learn, XGBoost
- Prophet and Darts for forecasting and time series modeling
Cloud & Ops
- Docker for containerization and portability
- AWS and Azure for scalable cloud architectures
- GitHub for version control and collaboration
- Hugging Face for model deployment and sharing
Computer Vision & Media Processing
- OpenCV for image and video processing
- MoviePy and FFmpeg for video automation
- Pillow for image manipulation
Generative AI
- LLMs (Claude, ChatGPT, Gemini) - Advanced proficiency
Professional Experience
ESPM (Superior School of Advertising and Marketing)
Data Scientist | Aug 2024 – Present
At ESPM, I work on restructuring the data lifecycle across extraction, loading, and transformation workflows, while also designing and maintaining a set of academic quality indicators that support strategic decision-making and pedagogical evaluation.
A major part of my role involves developing and structuring KPIs related to student attraction, retention, progression, and academic performance over time. These indicators allow us to monitor academic outcomes across different admission channels and understand how student performance evolves throughout the program.
In parallel, I help develop analytical products that support the segmentation and classification of newly admitted students, enabling early identification of those who may require academic reinforcement based on previously defined metrics and performance indicators.
Beyond the technical implementation, I actively contribute to the conceptual design of the metrics and frameworks used to map the full academic ecosystem. This includes translating educational goals into measurable indicators, validating assumptions with data, and communicating results to stakeholders.
I also work on complementary solutions such as dashboards that integrate machine learning outputs with academic metadata, as well as automation tools for academic operations. One example is the development of an internal exam-generation platform, which improves efficiency, reduces operational cost, and strengthens the data foundation required to support the KPIs built for institutional monitoring.
- Predictive Analytics: KPI dashboards for approval rates, retention, and academic completion
- Applied Research: ML/AI approaches applied to interdisciplinary academic datasets
- Data Products: Analytics solutions supporting academic management and learning strategies
- Process Optimization: Resource allocation analysis and operational efficiency improvements
Tech: Python, Power BI, SQL, PySpark, MLflow, Azure
BINAHKI
Senior Data Scientist | Nov 2023 – Aug 2024
At BINAHKI, I worked on smart monitoring products for multinational clients such as ZF Friedrichshafen, using IoT sensor data to monitor industrial machines and optimize production processes.
The main project involved monitoring industrial furnaces used in steel production through electrical voltage signals. The sensor data was stored in a cloud-based SQL environment integrated with AWS S3, and the pipeline included data ingestion, processing, and model-based detection of operational cycles.
Using Python and TensorFlow, I trained time series models to learn normal behavior patterns and identify key events such as the start and end of each production batch. Based on this cycle-level detection, I built statistical energy-consumption parameters that enabled forecasting of monthly usage and measurement of production efficiency.
In addition to furnace monitoring, similar IoT solutions were deployed to other industrial machines with the same objective: anomaly detection for predictive maintenance. These models supported early identification of abnormal behavior in equipment involving pumps, torque systems, and mechanical load.
Beyond product development, I also contributed to the design and architecture of data science solutions for public innovation calls. These initiatives resulted in successful proposals, including a project for a sanitation company focused on correcting missing and anomalous records using multivariate time series methods and statistical models such as ARIMAX.
Tech: Python, SQL, AWS S3, SageMaker
CBPF (Brazilian Center for Research in Physics)
Scientist | Nov 2017 – Jan 2024
My Master's research focused on combining systematic investigation with advanced laboratory methodology and quantitative interpretation of experimental results. This experience strengthened my ability to convert complex measurements into reliable scientific conclusions.
During my PhD, I conducted a systematic optimization study of physical properties under controlled experimental conditions. I compared the same material across multiple processing configurations and extracted insights from structural, microstructural, electronic, and magnetic datasets. This work required high analytical rigor, reproducibility, and careful interpretation of nonlinear behavior, reinforcing my ability to investigate complex systems and identify causal relationships in data.
This research path strengthened key Data Science capabilities such as experimental design thinking, statistical reasoning, pattern recognition, critical validation, and technical communication. Most importantly, it reinforced the ability to structure findings into clear narratives supported by evidence and consistent methodology.
Publications
2025
Enhanced oxygen mobility in NiAg alloy catalysts for methane dry reforming
ARG Caranton, AVP Lino, C Macchiutti, NRC Huaman, E Annese
Catalysis Today, 455, 115316 (2025)
DOI
2024
Optimization of the exchange bias effect in La₁.₅Sr₀.₅CoMnO₆
Calazans Macchiutti - Ph.D. Thesis, CBPF (2024)
PDF
Tuning the spontaneous exchange bias effect in La₁.₅Sr₀.₅CoMnO₆
JR Jesus, FB Carneiro, L Bufaiçal, RA Klein, Q Zhang, C Macchiutti
Physical Review Materials, 8(4), 044408 (2024)
DOI
2023
3d and 5d electronic structures in Ba- and Ca-doped double perovskites
JRL Mardegan, LSI Veiga, T Pohlmann, et al., C Macchiutti
Physical Review B, 107(21), 214427 (2023)
DOI
2022
Structural, electronic and magnetic properties of La₁.₅Ca₀.₅(Co₀.₅Fe₀.₅)IrO₆
L Bufaiçal, MAV Heringer, JR Jesus, A Caytuero, C Macchiutti, EM Bittar
J. Magn. Magn. Mater., 556, 169408 (2022)
DOI
2021
Master's Thesis: Synthesis and characterization of single crystals
Calazans Macchiutti - CBPF (2021)
PDF
Absence of zero-field-cooled exchange bias effect in single crystalline compounds
C Macchiutti, JR Jesus, FB Carneiro, L Bufaiçal, M Ciomaga Hatnean
Physical Review Materials, 5(9), 094402 (2021)
DOI
2020
Tuning the spontaneous exchange bias effect with Ba to Sr partial substitution
M Boldrin, AG Silva, LT Coutrim, JR Jesus, C Macchiutti, EM Bittar
Applied Physics Letters, 117(21), 212402 (2020)
DOI
Unveiling charge density wave quantum phase transitions by x-ray diffraction
FB Carneiro, LSI Veiga, JRL Mardegan, R Khan, C Macchiutti, A López
Physical Review B, 101(19), 195135 (2020)
DOI
2019
Zero-field-cooled exchange bias effect in phase-segregated La₂₋ₓAₓCoMnO₆₋δ
LT Coutrim, D Rigitano, C Macchiutti, TJA Mori, R Lora-Serrano
Physical Review B, 100(5), 054428 (2019)
DOI
2017
Física e Música (Physics and Music)
FN Grillo, HE Perez, C Macchiutti - Book, MNPEF (2017)
Co-author of Chapter 3 on physics and music.
PDF
Projects
Fraud Detection (KYC/KYT)
ML-based fraud detection comparing supervised and unsupervised approaches.
Python
XGBoost
TensorFlow
Scikit-learn
Logistic Data Assessment
Analytics suite for tracking delays, seasonality, and predictive risk scoring.
Python
Pandas
LightGBM
FastAPI
Streamlit Applications
Interactive web applications for data visualization and analysis.
Python
Streamlit
Plotly
Voltage Time Series LSTM
LSTM neural network for voltage prediction and anomaly detection.
Python
TensorFlow
LSTM
Magnetism Techniques Thesis
Research code and analysis for magnetism techniques thesis work.
Research
Data Analysis
GitHub
View all repositories on GitHub
Logistic Data Assessment
Exploratory and predictive analytics for operational efficiency in logistic networks
Screenshots
Departure Delays
Seasonality
Model Insights
Project Overview
This project consolidates exploratory analysis, KPI monitoring, and predictive modeling for a logistic operation handling thousands of daily deliveries. The objective is to surface actionable insights for dispatch planning, identify systemic bottlenecks, and forecast risk factors that affect service-level agreements.
Key Questions Answered
- What routes, hubs, and shifts are the main contributors to departure delays?
- How do seasonal patterns and workload peaks affect fleet punctuality?
- Which operational variables most influence delay risk according to the predictive models?
Analytical Highlights
- Interactive dashboards covering delay distribution, seasonal effects, and operational KPIs
- Machine learning pipeline (gradient boosting and logistic regression) for delay risk scoring
- Explainability layer with feature importance, SHAP interpretation, and scenario simulation
- Automated data quality checks and anomaly alerts for incoming logistic batches
Workflow
- Data ingestion from transactional systems and IoT telemetry with validation routines
- Feature engineering combining route metadata, fleet schedules, weather, and customer SLAs
- Model training and cross-validation with monitoring for drift and performance decay
- Packaging of results into dashboards and API endpoints for operational teams
Technologies
Python
Pandas
Polars
Scikit-learn
LightGBM
Plotly
FastAPI
Results & Impact
The combined analytical assets helped reduce average departure delays by 18% in pilot operations, while the predictive service highlighted upcoming high-risk routes with over 80% precision. Operational teams now rely on the dashboards to reallocate resources earlier and plan contingencies for peak seasons.
View on GitHub
← Back to Projects
Streamlit Applications
Interactive web applications built with Streamlit for data visualization and analysis
Screenshots
Main Dashboard
Analytics View
Visualizations
Project Overview
This project showcases a collection of interactive web applications built using Streamlit, a powerful Python framework for creating data applications. These applications demonstrate various aspects of data science, machine learning, and data visualization.
Features
- Interactive data visualization dashboards
- Machine learning model demonstrations
- Real-time data analysis tools
- User-friendly web interfaces
- Responsive design for various screen sizes
Applications Included
Data Visualization Dashboard
Interactive charts and graphs for exploring datasets with various filtering and customization options.
Machine Learning Playground
Hands-on interface for testing different ML algorithms and visualizing their performance.
Statistical Analysis Tool
Comprehensive statistical analysis with automated report generation and visualization.
Technologies
Python
Streamlit
Pandas
Plotly
NumPy
Scikit-learn
Getting Started
git clone https://github.com/Calazansmacchiutti/StreamlitApp.git
cd StreamlitApp
pip install -r requirements.txt
streamlit run app.py
View on GitHub
← Back to Projects
Voltage Time Series LSTM
LSTM neural network for voltage time series prediction and analysis
Screenshots
LSTM Model
Predictions
Training
Project Overview
This project implements a Long Short-Term Memory (LSTM) neural network for voltage time series prediction and analysis. The model is designed to forecast voltage patterns and detect anomalies in electrical systems using advanced deep learning techniques.
Features
- LSTM neural network architecture for time series forecasting
- Voltage pattern recognition and prediction
- Real-time anomaly detection capabilities
- Interactive data visualization and analysis tools
- Model performance evaluation and metrics
- Scalable architecture for different time series lengths
Key Components
Data Preprocessing
Comprehensive data cleaning, normalization, and feature engineering for voltage time series data.
LSTM Model Architecture
Multi-layer LSTM network with dropout regularization and optimized hyperparameters for voltage prediction.
Prediction Analysis
Advanced forecasting capabilities with confidence intervals and performance metrics evaluation.
Visualization Dashboard
Interactive plots showing original vs predicted values, loss curves, and model performance metrics.
Technical Details
Model Architecture
- Multi-layer LSTM with configurable hidden units
- Dropout layers for regularization
- Dense output layer for regression
- Adam optimizer with adaptive learning rate
Data Processing
- Time series windowing for sequence generation
- Min-Max normalization for stable training
- Train/validation/test split with temporal ordering
- Feature scaling and inverse transformation
Technologies
Python
TensorFlow
Keras
NumPy
Pandas
Matplotlib
LSTM
Getting Started
git clone https://github.com/Calazansmacchiutti/voltage-times-series-lstm.git
cd voltage-times-series-lstm
pip install -r requirements.txt
python main.py
View on GitHub
← Back to Projects
Fraud Detection (KYC/KYT)
Data Science approaches for detecting fraudulent transactions using supervised and unsupervised learning
Screenshots
📄 Research Paper (click to download)
PCA Features
Model Comparison
Amount Analysis
Project Overview
This project implements a comprehensive fraud detection system combining Know Your Customer (KYC) and Know Your Transaction (KYT) methodologies. Using the Credit Card Fraud Detection dataset, the system evaluates multiple machine learning approaches including supervised classifiers, unsupervised anomaly detection, and deep learning autoencoders to identify fraudulent transactions in highly imbalanced datasets.
Key Questions Answered
- How do supervised models (XGBoost, Random Forest) compare against unsupervised approaches?
- Which PCA-transformed features show the strongest discriminative power?
- What is the optimal threshold strategy for extreme class imbalance (~0.17% fraud rate)?
- How can autoencoder reconstruction error be leveraged for anomaly detection?
Analytical Highlights
- XGBoost achieved highest Average Precision (0.858) followed by Random Forest (0.826)
- ROC-AUC scores above 0.96 for all models, with XGBoost reaching 0.969
- Autoencoder-based anomaly detection with tuned threshold achieving AP of 0.510
- Feature importance analysis revealing V14, V4, V12, V10 as top discriminators
Technologies
Python
Scikit-learn
XGBoost
TensorFlow/Keras
Pandas
Matplotlib
Results & Impact
Supervised learning significantly outperforms unsupervised methods when labeled data is available. XGBoost emerged as the best performer with 85.8% Average Precision, detecting 86% of fraudulent transactions while maintaining 51% precision. The autoencoder approach provides a viable alternative for cold-start scenarios when labeled fraud examples are scarce.
Download Paper
View on GitHub
← Back to Projects
Curriculum Vitae
Education
Ph.D. in Physics
CBPF | 2020-2024
"Optimization of the exchange bias effect in La₁.₅Sr₀.₅CoMnO₆"
Master's in Physics
CBPF | 2018-2021
"Synthesis and characterization of single crystals of La₀.₅(Ca,Sr)₀.₅CoMnO₆"
Experience
Senior Data Scientist - ESPM
Aug 2024 - Present
Predictive modeling, scientific partnerships, AI innovation in education.
Senior Data Scientist - BINAHKI
Nov 2023 - Aug 2024
ML models for industrial applications, smart monitoring, automation pipelines.
Skills
Programming
- Python
- R
- SQL
- MATLAB
- FastAPI
- Pandas/Polars
Data Science
- Machine Learning
- Hugging Face
- Statistics
- Causal Inference
- Visualization
- Time Series
Engineering
- System Design
- Project Mgmt
Research
- Condensed Matter
- Materials Science
- X-ray Diffraction
- Magnetism
Links
Lattes
Scholar
ResearchGate
LinkedIn