Data Science · Melbourne

Hi, I'm Sam.

I'm a data scientist and current applied mathematics student based in Melbourne. My work largely focuses on statistical modelling and machine learning, and academic interests of mine span stochastic processes, software engineering, linguistics, and beyond.

I currently work in data science at MedHealth in Melbourne alongisde my degree, where I enjoy regular opportunities to take initiative to work independently to research, build, and validate models, as well as to participate in team environments where I am able to learn from experts with diverse expertise.

Beyond work and uni, I love to read, play cricket and golf, and go camping.

Headshot of Sam Murphy

I wrote my first line of Python code while I was in primary school, and ever since have been captivated by new ways to approach problems computationally. Through my later years of high school my love for maths truly flourished and that's how I eventually led myself from software development towards data science: a new and ever-changing paradigm which allows me to apply my skillset creatively.

Particularly I am excited by problems involving uncertain outcomes, and opportunities to dig into large datasets using the tools at my disposal. I like taking initiative in the projects I undertake, but also enjoy collaboration in data science pursuits, sharing ideas and gaining new understandings from experienced professionals with a variety of interests.

Modelling

Statistical learning · classification · regression · survival analysis · time series · model validation

Data

Python · SQL · pandas · Spark · Databricks

Scientific computing

NumPy · SciPy · statsmodels · numerical methods · stochastic simulation

Software

Matlab · Git · C++ · APIs ·

Current

Data Science Intern

MedHealth, Melbourne (Dec 2025–Present)

Undergraduate

BSc (Applied Mathematics)

University of Melbourne, Parkville (2024–2026)

Modelling experiments

Interactive investigations into statistical modelling, stochastic systems and quantitative methods. Each experiment focuses on a result with a practical implication for how models are built, validated or used.

A short selection of professional and independent work that I'm particularly fond of.

01

Applied data science

Return-to-work risk modelling

2026

Developed statistical models to estimate return to work probability for different KPI thresholds using longitudinal case and operational data, serving clients across a broad range of industries and disciplines. Rigorously validated the model out-of-sample and built a triage prioritisation engine based on the insights from the model.

Problem

Estimating probability of extended absence or non-return to work

Methods

Classification · survival analysis · calibration · temporal validation

Stack

Python · Spark · Databricks

Professional work. Technical description excludes confidential data and implementation details.

02

Quantitative research

Cross-asset Yield and Commodity Signals

2025

Systematic research pipeline to test whether information in the US treasury curve and commodity markets offer predictive utility about forward returns across equities, fixed income, commodities, and FX.

Constructs rolling yield curve, momentum, and mean-reversion signals and screens candidate predictors across multiple time horizons using statistical analysis and converts selected forecasts into bounded trading positions. A final hold-out backtest is then run.

Research

Yield curve spreads · Commodities markets · Relative-value signals · Financial modelling

Validation

Walk-forward refit · transaction cost sensitivity · alpha regression

Stack

Python · numpy · statsmodels · scipy · FRED · Yahoo Finance

03

Histoinformatics / Computational Linguistics

Pursuing Linear A via Hidden Markov Models

2026

Inspired by previous work led by Dr Brent Davis, I am researching the use of Discrete Time Markov Chains and Hidden Markov Models in attempting to decipher scripts and whether or not this could be applied to ancient Minoan Linear A.

This project is ongoing to mixed results as one might expect, but it is always exciting to be greeted by an opportunity to work at the intersection of my professional and extracurricular interests.

Research

Historical linguistics · Computational linguistics · Translation of ancient scripts

Core

Model selection · Corpus encoding · Statistical testing

Stack

Python, C++

04

Quantitative Finance

Exotic Options Pricing Engine

2025

Built a modular derivatives pricing engine for barrier, multi-asset, path-dependent, and early-exercise derivative securities. The project compares finite differencing PDE methods, Monte Carlo simulation, variance reduction techniques, and other numerical mehtods according to the structures and dimensionalities of each modelled product.

The engine prices Asian, barrier, American, and basket derivatives, computes numerical option greeks, and produces confidence intervals.

Methods

Finite differences · Monte Carlo · Variance Reduction · Convergence Analysis

Products

Vanilla and digital options · Asian options · American options · Path-dependent payoffs

Stack

Python · numpy · numerical linear algebra · OOP

05

Applied Data Science

Horse Racing Prediction Engine

2024–Present

Built and continue to develop a data-driven system for pricing thoroughbred horse races in Australia that collects and standardises races, runner and market data, and produces probability estimates. I am then able to discretionarily compare these modelled probabilities against available market odds to identify potential value in the form of mispricings.

NB: my interest in horse racing is primarily from a modelling and market-pricing perspective rather than as an avid bettor. I do not endorse the significant animal welfare and broader ethical issues associated with the industry, and any use of betting markets should be approached responsibly.

Methods

Web scraping · Automated data collection · Feature engineering · Model evaluation

Analysis

Market-implied probabilities · Model prices · Calibration and performance metrics

Stack

Python · numpy · APIs

06

Personal Project

Bookcase: A Private Home for Every Book She Loves

2026

My girlfriend loves to read, as well as to buy books she may one day read but can't quite get around to yet. With many hundreds of books to keep track of she was starting to get a bit tired of having to, let's say, O(n) search (she ruled out my alphabetical bookshelf system long ago) her entire bookshelf to find if she already owns a title before going shopping. Behold, Bookcase, a simple website I built while she was away that allows her to file every book on her shelf, every book she would like to put on her shelf, and every book she has read but has not quite made it to the shelf.

It gives you real-time reading progress statistics for series, top genre information, periodic reading stats such as "You've read 10 books this year", ratings, wish lists, etc., and is an ongoing project I continue to build with her guidance on what she would like it to do.

The next stage I am currently working on for this project is implementing a recommendation engine that, based on a blend of long-term and short-term historical reading habits, identifies what books she has not read that she might like.

Methods

Database development ·

The next step

Machine Learning book recommendation engine

Stack

NextJS · Vercel · Postgres

Get in touch.

I'm always up for a chat; whether it be about data science, mathematics, and related opportunities (or even the upcoming summer of Test cricket!). Find me at any of the below.