About Skills Projects Experience Contact
Open to AI/ML opportunities

Mark Oliver
Rodrigues

AI/ML Engineer

Building intelligent systems that transform data into real-world impact. Specializing in RAG systems, semantic search, and LLM-powered applications.

AI / ML

Computer Engineering graduate (Honours in AI/ML) and a working AI/ML Engineer with hands-on experience building production-grade RAG systems, semantic search pipelines, and LLM-powered applications. I've deployed real projects using vector databases, transformer models, and cloud infrastructure.

I'm passionate about building intelligent, scalable applications — from AI Agents and RAG systems to semantic search pipelines and NLP projects. I enjoy transforming data into real-world impact and continuously learning modern AI tooling.

I graduated in Computer Engineering with Honours in AI/ML from St. Francis Institute of Technology and currently work as a Junior AI Engineer at ComplAIBridge, shipping document-AI and compliance backends. Already solving real business problems in production, I'm now looking for better opportunities to take on bigger challenges and grow my impact at scale.

21+
GitHub Repos
4
Vector DBs Used
15+
AI/ML Projects
7+
Models Deployed
Programming & Tools
PythonSQLJavaScript MySQLPostgreSQLGit GitHubDockerTableauPower BI
Frameworks & Libraries
PyTorchTensorFlowKeras Scikit-learnFastAPIFlask LangChainPandasNumPy StreamlitPydanticSQLAlchemy pytestMatplotlibSeabornPlotly
AI & Deep Learning
NLPRAGLLMs TransformersCLIPRoBERTa GNNRegressionClassification Random ForestsXGBoost LightGBMCatBoostSHAP LIMEOptunaRDKit ANNCNNRNN OpenCVOCRMediaPipe
Vector Databases
FAISSMilvusQdrant ChromaDB
MLOps & Deployment
DockerDVCMLflow DaskProphetCI/CD Hugging Face SpacesAzure AI FoundryJWT Auth
My Projects — from github.com/Mark007-R
AI-Customer-Ops-Engine

AI infrastructure that gives customer-service agents in regulated industries a long-term memory and a safe, auditable way to make decisions — kept isolated per client, with a rule-based approval queue for sensitive actions. 543 tests passing.

PythonFastAPIPostgresSQLAlchemypytestDocker
543 tests passing
Multi-tenant isolation
Idempotent ingestion
Fraud-Detection-MLOps

A fraud-detection system for card payments, built as a full MLOps pipeline — it trains an XGBoost model, watches live data for drift, and automatically retrains itself when the patterns shift. Comes with a live dashboard for monitoring.

PythonXGBoostDaskDVCMLflowFastAPIPostgresStreamlitOptuna
Drift detect → promote in ~18s
4ms alias-flip rollback
Bit-exact Pandas/Dask
Semantic-Movie-Recommender

A movie recommender that searches 9,000+ films by meaning — describe a plot or upload a poster and it finds similar movies in under 100ms, matching on both story and visuals.

PythonStreamlitSentence TransformersOpenAI CLIPMilvusPyTorch
9,000+ movies indexed
<100ms latency
Multi-modal search
Diagram-Structure-Extractor

Reads architecture diagrams for you — finds the boxes, labels, and arrows automatically, then exports the whole structure as annotated images and clean JSON/CSV.

PythonOpenCVOCRComputer Vision
Annotated outputs
Relationship graphs
JSON/CSV export
Restaurant-Intelligence-Platform

Reads 10,000+ restaurant reviews and turns them into insights — gauging sentiment, sorting complaints into categories, and letting you chat with a bot that answers questions about the reviews in under a second.

PythonFlaskFAISSSentenceTransformersVADERPandas
10,000+ reviews processed
Sub-second response
5+ categories
Document-QA-RAG

Upload a PDF and ask questions about it — the app finds the relevant passages and gives you grounded, source-backed answers instead of guesses.

PythonFlaskFAISSSentenceTransformersLLaMA
PDF document QA
Multi-LLM support
RAG pipeline
Stock-Price-Forecaster

Predicts stock prices with an LSTM model and pairs the forecast with live news and sentiment, so you get the full picture before making an investment decision.

PythonStreamlitFlaskLSTMTensorFlowNLP
LSTM predictions
Live news integration
Sentiment analysis
Code-Review-Agent

An AI agent that reviews your code — it spots bugs, security holes, and slow spots, then hands back fixes with severity ratings and the exact lines to change.

PythonAI AgentsCode AnalysisJSON
Automated code review
Security scanning
Structured findings
AI-Personal-Finance-Manager

A personal-finance app that scans your receipts, auto-sorts your spending into categories, flags unusual charges and forgotten subscriptions, and forecasts your cash flow — plus risk-profiled investment tips. Every feature is benchmarked against a frontier LLM and served from a secure, multi-tenant FastAPI backend.

PythonFastAPIscikit-learnProphetOCRJWT
Receipt OCR + categorizer
Anomaly + forecast
66 tests passing
Open-Source Contributions — shipped phases on a 5-project YC-style portfolio
Fraud Detection System
Contributor

An open-source card-fraud detector I shipped Phases 1–7 of — feature engineering, model face-offs, and a head-to-head against a frontier LLM, ending with a CatBoost model that hit F1 = 1.000 while running ~240,000× faster and ~45,000× cheaper.

PythonCatBoostXGBoostLightGBMFastAPIDockerSHAPStreamlit
CatBoost F1 = 1.000
94 regression tests
7 merged PRs
Visual Product Search Engine
Contributor

Search fashion products by image — I built the visual-search path using CLIP embeddings plus a color histogram, tuned over 300 trials. The surprise finding: a tiny 48D color feature beat a 2048D ResNet50.

PythonCLIPFAISSFastAPIDockerStreamlitOpenCVOptuna
R@1 = 0.683
Sub-ms search
12 visual-only configs
Drug Molecule Property Prediction
Contributor

Predicts a drug molecule's properties straight from its structure — I shipped Phases 4–7, including a 3-model ensemble that matches a graph neural network and a SHAP-vs-LIME study of why the model makes each call.

PythonRDKitGNNOptunaSHAPLIMEStreamlit
Ensemble matches GNN+Edge
50/50 tests passing
SHAP + LIME comparison
Healthcare Readmission Predictor
Contributor

Predicts which patients are likely to be readmitted to hospital — I shipped the dataset and EDA, the feature engineering, per-subgroup risk thresholds, and a SHAP deep dive into what actually drives the risk.

PythonXGBoostSHAPPandasScikit-learn
Subgroup-threshold routing
53% SHAP from 2 features
4 merged PRs
Legal Contract Analyzer
Contributor

Classifies the clauses inside legal contracts — I shipped Phases 3–7, where a tuned LightGBM blend reached macro-F1 0.72, beating both RoBERTa-large and a frontier LLM.

PythonLightGBMRoBERTascikit-learnOptuna
Macro-F1 = 0.7163
Beats RoBERTa by +0.066
64/64 tests
Junior AI Engineer Full-time
ComplAIBridge
Apr 2026 -- Present
  • Built five production AI services powering ComplAIBridge's compliance platform
  • Shipped a tool-calling project chatbot — Mistral, FAISS + MongoDB retrieval, prompt-injection/PII/content-safety guardrails, SSE streaming, and source-cited answers
  • Shipped a chatbot-audit API that live-probes bots for jailbreak, PII leakage, toxicity, and bias, scoring them against ISO 42001, NIST AI RMF, EU AI Act, SOC 2, HIPAA, and GDPR
  • Built an MCP-enabled compliance engine that scans GitHub PRs, commits, prompts, and full repos against 31 control catalogs with per-line AI-session attribution, webhooks, and async jobs — plus auto-generated architecture diagrams with design-change impact analysis
  • Delivered an LLM remittance/bank-statement parser (Docling + PaddleOCR) with hybrid three-way reconciliation
  • Delivered a news-grounded PESTLE risk scan/monitor engine, an agentic SOP control-verification service probing live AWS/Azure/GitHub/Jira systems, and Jira-to-requirements/risks/controls and architecture-design analysis APIs
AI Data Trainer Contract
Innodata Inc.
Nov 2025 -- Jan 2026
  • Annotated and labeled large-scale text and image datasets for supervised model training
  • Performed LLM instruction tuning — writing prompts, editing generated data, and refining model responses
  • Labeled computer-vision datasets on human attributes and visual elements to strict annotation and quality guidelines
  • Supported model evaluation and QA by reviewing AI outputs, correcting errors, and providing structured feedback used to improve performance
  • Worked within a production pipeline requiring high accuracy, confidentiality, and fast turnaround
B.E. Computer Engineering
Honours in AI/ML -- St. Francis Institute of Technology, Borivli
Graduated May 2026
CGPA: 7.85 / 10
HSC (Computer Science)
Shri T.P. Bhatia College, Kandivli
June 2022
Generative AI: Working with Large Language Models
LinkedIn Learning
Advanced SQL
Kaggle
Data Analytics Job Simulation
Deloitte (Forage)

I'm currently looking for better opportunities in AI/ML engineering. Whether you have a question, a project idea, or just want to say hi, I'd love to hear from you.

markrodrigues2689@gmail.com
Mumbai, India