About Me

Calazans Macchiutti

Welcome to my website! I'm Calazans Macchiutti, a Data Scientist with a Ph.D. in Physics (CBPF) and a strong background in statistical modeling, scientific reasoning, and applied machine learning.

Today, I work as a Data Scientist at ESPM, where I build data solutions that support strategic academic decision-making, including predictive analytics, KPI development, automation pipelines, and dashboards. I enjoy turning complex and messy data into clear insights, reliable metrics, and scalable products that improve real-world processes.

My academic training gave me a rigorous foundation in problem-solving, experimental design, and quantitative analysis. Over time, I transitioned that expertise into industry-focused work, applying machine learning to time series, anomaly detection, and performance forecasting — always aiming for practical impact, clarity, and measurable results.

Areas of Focus

  • Machine Learning & Predictive Modeling: classification, forecasting, anomaly detection
  • Time Series Analytics: pattern recognition, monitoring, efficiency indicators
  • Applied Data Science: dashboards, KPI frameworks, data quality metrics
  • Automation & Data Pipelines: scalable workflows for analytics and reporting
  • AI for Decision Support: insights that improve processes and operations

Professional Experience

  • Data Scientist — ESPM (2024–Present)
  • Senior Data Scientist — BINAHKI (2023–2024)
  • Ph.D. Researcher — CBPF (2020–2024)

I'm especially motivated by projects where data can drive better decisions, optimize resources, and improve outcomes. If you're looking for a Data Scientist who combines technical depth, strong analytical thinking, and a practical product-driven mindset, feel free to reach out.

Academic Profile

Academic Profile

Research in Condensed Matter Physics and Materials Science

Research Focus

During my doctoral program, I focused my studies on the physical properties of transition metal oxides, with a special emphasis on the investigation of double perovskite-type structures and their impact on the "exchange bias" phenomenon.

My research covers a variety of techniques:

  • Magnetic measurements
  • Electrical and thermal transport studies
  • X-ray diffraction
  • Magnetoresistance analysis
  • Investigation of strong electronic correlation phenomena

Materials Synthesis Experience

Extensive experience in materials synthesis through various methods:

  • Sol-gel synthesis
  • Co-precipitation methods
  • Solid-state synthesis
  • Floating zone techniques

All syntheses included both polycrystals and single crystals of ternary and quaternary rare earth oxides, as well as binary single crystals such as V₂O₅.

Characterization Techniques

  • X-ray photoelectron spectroscopy (XPS)
  • Raman spectroscopy
  • Atomic force microscopy (AFM)

Early Research

  • Musical Acoustics and Environmental Acoustics
  • Applications of electromagnetic waves in magnetic resonance
  • Analysis of sound waves in music

View all publications →

Data Science

Data Scientist

Professional Summary

My journey has been shaped by the development and application of statistical models, with a strong focus on identifying patterns, understanding trends, and generating actionable insights.

I developed a strong foundation in scientific problem-solving, quantitative modeling, and rigorous analytical reasoning through end-to-end research projects. This experience includes hypothesis formulation, experimental planning, data acquisition, and the interpretation of complex multivariate phenomena. With this background, I present below a detailed overview of my professional trajectory from a technical perspective.

Technical Skills

Programming & Data Manipulation

  • Python (Pandas, Polars, NumPy)
  • SQL for querying, transformation, and data management
  • Dask and PySpark for large-scale processing

Data Visualization & Statistical Analysis

  • Matplotlib, Seaborn, Plotly
  • SciPy and Statsmodels for statistical modeling and inference

Machine Learning & AI

  • TensorFlow, Keras, Scikit-learn, XGBoost
  • Prophet and Darts for forecasting and time series modeling

Cloud & Ops

  • Docker for containerization and portability
  • AWS and Azure for scalable cloud architectures
  • GitHub for version control and collaboration
  • Hugging Face for model deployment and sharing

Computer Vision & Media Processing

  • OpenCV for image and video processing
  • MoviePy and FFmpeg for video automation
  • Pillow for image manipulation

Generative AI

  • LLMs (Claude, ChatGPT, Gemini) - Advanced proficiency

Professional Experience

ESPM (Superior School of Advertising and Marketing)

Data Scientist | Aug 2024 – Present

At ESPM, I work on restructuring the data lifecycle across extraction, loading, and transformation workflows, while also designing and maintaining a set of academic quality indicators that support strategic decision-making and pedagogical evaluation.

A major part of my role involves developing and structuring KPIs related to student attraction, retention, progression, and academic performance over time. These indicators allow us to monitor academic outcomes across different admission channels and understand how student performance evolves throughout the program.

In parallel, I help develop analytical products that support the segmentation and classification of newly admitted students, enabling early identification of those who may require academic reinforcement based on previously defined metrics and performance indicators.

Beyond the technical implementation, I actively contribute to the conceptual design of the metrics and frameworks used to map the full academic ecosystem. This includes translating educational goals into measurable indicators, validating assumptions with data, and communicating results to stakeholders.

I also work on complementary solutions such as dashboards that integrate machine learning outputs with academic metadata, as well as automation tools for academic operations. One example is the development of an internal exam-generation platform, which improves efficiency, reduces operational cost, and strengthens the data foundation required to support the KPIs built for institutional monitoring.

  • Predictive Analytics: KPI dashboards for approval rates, retention, and academic completion
  • Applied Research: ML/AI approaches applied to interdisciplinary academic datasets
  • Data Products: Analytics solutions supporting academic management and learning strategies
  • Process Optimization: Resource allocation analysis and operational efficiency improvements

Tech: Python, Power BI, SQL, PySpark, MLflow, Azure

BINAHKI

Senior Data Scientist | Nov 2023 – Aug 2024

At BINAHKI, I worked on smart monitoring products for multinational clients such as ZF Friedrichshafen, using IoT sensor data to monitor industrial machines and optimize production processes.

The main project involved monitoring industrial furnaces used in steel production through electrical voltage signals. The sensor data was stored in a cloud-based SQL environment integrated with AWS S3, and the pipeline included data ingestion, processing, and model-based detection of operational cycles.

Using Python and TensorFlow, I trained time series models to learn normal behavior patterns and identify key events such as the start and end of each production batch. Based on this cycle-level detection, I built statistical energy-consumption parameters that enabled forecasting of monthly usage and measurement of production efficiency.

In addition to furnace monitoring, similar IoT solutions were deployed to other industrial machines with the same objective: anomaly detection for predictive maintenance. These models supported early identification of abnormal behavior in equipment involving pumps, torque systems, and mechanical load.

Beyond product development, I also contributed to the design and architecture of data science solutions for public innovation calls. These initiatives resulted in successful proposals, including a project for a sanitation company focused on correcting missing and anomalous records using multivariate time series methods and statistical models such as ARIMAX.

Tech: Python, SQL, AWS S3, SageMaker

CBPF (Brazilian Center for Research in Physics)

Scientist | Nov 2017 – Jan 2024

My Master's research focused on combining systematic investigation with advanced laboratory methodology and quantitative interpretation of experimental results. This experience strengthened my ability to convert complex measurements into reliable scientific conclusions.

During my PhD, I conducted a systematic optimization study of physical properties under controlled experimental conditions. I compared the same material across multiple processing configurations and extracted insights from structural, microstructural, electronic, and magnetic datasets. This work required high analytical rigor, reproducibility, and careful interpretation of nonlinear behavior, reinforcing my ability to investigate complex systems and identify causal relationships in data.

This research path strengthened key Data Science capabilities such as experimental design thinking, statistical reasoning, pattern recognition, critical validation, and technical communication. Most importantly, it reinforced the ability to structure findings into clear narratives supported by evidence and consistent methodology.

Publications

2025

Enhanced oxygen mobility in NiAg alloy catalysts for methane dry reforming

ARG Caranton, AVP Lino, C Macchiutti, NRC Huaman, E Annese

Catalysis Today, 455, 115316 (2025)

DOI

2024

Optimization of the exchange bias effect in La₁.₅Sr₀.₅CoMnO₆

Calazans Macchiutti - Ph.D. Thesis, CBPF (2024)

PDF

Tuning the spontaneous exchange bias effect in La₁.₅Sr₀.₅CoMnO₆

JR Jesus, FB Carneiro, L Bufaiçal, RA Klein, Q Zhang, C Macchiutti

Physical Review Materials, 8(4), 044408 (2024)

DOI

2023

3d and 5d electronic structures in Ba- and Ca-doped double perovskites

JRL Mardegan, LSI Veiga, T Pohlmann, et al., C Macchiutti

Physical Review B, 107(21), 214427 (2023)

DOI

2022

Structural, electronic and magnetic properties of La₁.₅Ca₀.₅(Co₀.₅Fe₀.₅)IrO₆

L Bufaiçal, MAV Heringer, JR Jesus, A Caytuero, C Macchiutti, EM Bittar

J. Magn. Magn. Mater., 556, 169408 (2022)

DOI

2021

Master's Thesis: Synthesis and characterization of single crystals

Calazans Macchiutti - CBPF (2021)

PDF

Absence of zero-field-cooled exchange bias effect in single crystalline compounds

C Macchiutti, JR Jesus, FB Carneiro, L Bufaiçal, M Ciomaga Hatnean

Physical Review Materials, 5(9), 094402 (2021)

DOI

2020

Tuning the spontaneous exchange bias effect with Ba to Sr partial substitution

M Boldrin, AG Silva, LT Coutrim, JR Jesus, C Macchiutti, EM Bittar

Applied Physics Letters, 117(21), 212402 (2020)

DOI

Unveiling charge density wave quantum phase transitions by x-ray diffraction

FB Carneiro, LSI Veiga, JRL Mardegan, R Khan, C Macchiutti, A López

Physical Review B, 101(19), 195135 (2020)

DOI

2019

Zero-field-cooled exchange bias effect in phase-segregated La₂₋ₓAₓCoMnO₆₋δ

LT Coutrim, D Rigitano, C Macchiutti, TJA Mori, R Lora-Serrano

Physical Review B, 100(5), 054428 (2019)

DOI

2017

Física e Música (Physics and Music)

FN Grillo, HE Perez, C Macchiutti - Book, MNPEF (2017)

Co-author of Chapter 3 on physics and music.

PDF

ADS Library | ORCID: 0000-0002-8571-8876

Projects

Fraud Detection (KYC/KYT)

ML-based fraud detection comparing supervised and unsupervised approaches.

Python XGBoost TensorFlow Scikit-learn

Logistic Data Assessment

Analytics suite for tracking delays, seasonality, and predictive risk scoring.

Python Pandas LightGBM FastAPI

Streamlit Applications

Interactive web applications for data visualization and analysis.

Python Streamlit Plotly

Voltage Time Series LSTM

LSTM neural network for voltage prediction and anomaly detection.

Python TensorFlow LSTM

Magnetism Techniques Thesis

Research code and analysis for magnetism techniques thesis work.

Research Data Analysis
GitHub

View all repositories on GitHub

Logistic Data Assessment

Exploratory and predictive analytics for operational efficiency in logistic networks

Screenshots

Project Overview

This project consolidates exploratory analysis, KPI monitoring, and predictive modeling for a logistic operation handling thousands of daily deliveries. The objective is to surface actionable insights for dispatch planning, identify systemic bottlenecks, and forecast risk factors that affect service-level agreements.

Key Questions Answered

  • What routes, hubs, and shifts are the main contributors to departure delays?
  • How do seasonal patterns and workload peaks affect fleet punctuality?
  • Which operational variables most influence delay risk according to the predictive models?

Analytical Highlights

  • Interactive dashboards covering delay distribution, seasonal effects, and operational KPIs
  • Machine learning pipeline (gradient boosting and logistic regression) for delay risk scoring
  • Explainability layer with feature importance, SHAP interpretation, and scenario simulation
  • Automated data quality checks and anomaly alerts for incoming logistic batches

Workflow

  1. Data ingestion from transactional systems and IoT telemetry with validation routines
  2. Feature engineering combining route metadata, fleet schedules, weather, and customer SLAs
  3. Model training and cross-validation with monitoring for drift and performance decay
  4. Packaging of results into dashboards and API endpoints for operational teams

Technologies

Python Pandas Polars Scikit-learn LightGBM Plotly FastAPI

Results & Impact

The combined analytical assets helped reduce average departure delays by 18% in pilot operations, while the predictive service highlighted upcoming high-risk routes with over 80% precision. Operational teams now rely on the dashboards to reallocate resources earlier and plan contingencies for peak seasons.

View on GitHub ← Back to Projects

Streamlit Applications

Interactive web applications built with Streamlit for data visualization and analysis

Screenshots

Project Overview

This project showcases a collection of interactive web applications built using Streamlit, a powerful Python framework for creating data applications. These applications demonstrate various aspects of data science, machine learning, and data visualization.

Features

  • Interactive data visualization dashboards
  • Machine learning model demonstrations
  • Real-time data analysis tools
  • User-friendly web interfaces
  • Responsive design for various screen sizes

Applications Included

Data Visualization Dashboard

Interactive charts and graphs for exploring datasets with various filtering and customization options.

Machine Learning Playground

Hands-on interface for testing different ML algorithms and visualizing their performance.

Statistical Analysis Tool

Comprehensive statistical analysis with automated report generation and visualization.

Technologies

Python Streamlit Pandas Plotly NumPy Scikit-learn

Getting Started

git clone https://github.com/Calazansmacchiutti/StreamlitApp.git cd StreamlitApp pip install -r requirements.txt streamlit run app.py

View on GitHub ← Back to Projects

Voltage Time Series LSTM

LSTM neural network for voltage time series prediction and analysis

Screenshots

Project Overview

This project implements a Long Short-Term Memory (LSTM) neural network for voltage time series prediction and analysis. The model is designed to forecast voltage patterns and detect anomalies in electrical systems using advanced deep learning techniques.

Features

  • LSTM neural network architecture for time series forecasting
  • Voltage pattern recognition and prediction
  • Real-time anomaly detection capabilities
  • Interactive data visualization and analysis tools
  • Model performance evaluation and metrics
  • Scalable architecture for different time series lengths

Key Components

Data Preprocessing

Comprehensive data cleaning, normalization, and feature engineering for voltage time series data.

LSTM Model Architecture

Multi-layer LSTM network with dropout regularization and optimized hyperparameters for voltage prediction.

Prediction Analysis

Advanced forecasting capabilities with confidence intervals and performance metrics evaluation.

Visualization Dashboard

Interactive plots showing original vs predicted values, loss curves, and model performance metrics.

Technical Details

Model Architecture

  • Multi-layer LSTM with configurable hidden units
  • Dropout layers for regularization
  • Dense output layer for regression
  • Adam optimizer with adaptive learning rate

Data Processing

  • Time series windowing for sequence generation
  • Min-Max normalization for stable training
  • Train/validation/test split with temporal ordering
  • Feature scaling and inverse transformation

Technologies

Python TensorFlow Keras NumPy Pandas Matplotlib LSTM

Getting Started

git clone https://github.com/Calazansmacchiutti/voltage-times-series-lstm.git cd voltage-times-series-lstm pip install -r requirements.txt python main.py

View on GitHub ← Back to Projects

Fraud Detection (KYC/KYT)

Data Science approaches for detecting fraudulent transactions using supervised and unsupervised learning

Screenshots

Project Overview

This project implements a comprehensive fraud detection system combining Know Your Customer (KYC) and Know Your Transaction (KYT) methodologies. Using the Credit Card Fraud Detection dataset, the system evaluates multiple machine learning approaches including supervised classifiers, unsupervised anomaly detection, and deep learning autoencoders to identify fraudulent transactions in highly imbalanced datasets.

Key Questions Answered

  • How do supervised models (XGBoost, Random Forest) compare against unsupervised approaches?
  • Which PCA-transformed features show the strongest discriminative power?
  • What is the optimal threshold strategy for extreme class imbalance (~0.17% fraud rate)?
  • How can autoencoder reconstruction error be leveraged for anomaly detection?

Analytical Highlights

  • XGBoost achieved highest Average Precision (0.858) followed by Random Forest (0.826)
  • ROC-AUC scores above 0.96 for all models, with XGBoost reaching 0.969
  • Autoencoder-based anomaly detection with tuned threshold achieving AP of 0.510
  • Feature importance analysis revealing V14, V4, V12, V10 as top discriminators

Technologies

Python Scikit-learn XGBoost TensorFlow/Keras Pandas Matplotlib

Results & Impact

Supervised learning significantly outperforms unsupervised methods when labeled data is available. XGBoost emerged as the best performer with 85.8% Average Precision, detecting 86% of fraudulent transactions while maintaining 51% precision. The autoencoder approach provides a viable alternative for cold-start scenarios when labeled fraud examples are scarce.

Download Paper View on GitHub ← Back to Projects

Curriculum Vitae

Education

Ph.D. in Physics

CBPF | 2020-2024
"Optimization of the exchange bias effect in La₁.₅Sr₀.₅CoMnO₆"

Master's in Physics

CBPF | 2018-2021
"Synthesis and characterization of single crystals of La₀.₅(Ca,Sr)₀.₅CoMnO₆"

Experience

Senior Data Scientist - ESPM

Aug 2024 - Present
Predictive modeling, scientific partnerships, AI innovation in education.

Senior Data Scientist - BINAHKI

Nov 2023 - Aug 2024
ML models for industrial applications, smart monitoring, automation pipelines.

Skills

Programming

  • Python
  • R
  • SQL
  • MATLAB
  • FastAPI
  • Pandas/Polars

Data Science

  • Machine Learning
  • Hugging Face
  • Statistics
  • Causal Inference
  • Visualization
  • Time Series

Engineering

  • System Design
  • Project Mgmt

Research

  • Condensed Matter
  • Materials Science
  • X-ray Diffraction
  • Magnetism

Links

Lattes Scholar ResearchGate LinkedIn