↓ Skip to main content

Spot the Pain

Table of Contents

The Problem
#

Automated pain assessment can support healthcare and rehabilitation by making pain measurement more objective. Most research relies on facial expressions, while body movement, which people also use to express and protect themselves from pain, has received far less attention.

My Master’s thesis at Linnaeus University investigated whether skeleton pose data could be used on its own, or combined with facial features, for three tasks:

  1. Pain recognition: is the person in pain?
  2. Intensity estimation: how strong is the pain?
  3. Area classification: where is the pain located?

Approach
#

I extracted two modalities from video recordings of people performing an overhead deep squat:

  • Body: 17 skeleton keypoints per frame with PoseNet
  • Face: facial action units with OpenFace, based on the Prkachin and Solomon Pain Intensity (PSPI) scale

I trained two architectures suited to movement over time, a hybrid CNN-BiLSTM and a recurrent CNN (RCNN). I compared body-only models with three ways of combining (fusing) the two modalities:

graph TD
    A[Video] --> B[PoseNet: 17 body keypoints]
    A --> C[OpenFace: facial action units]
    B --> D{Fusion strategy}
    C --> D
    D -->|Early| E[Combine inputs]
    D -->|Late| F[Average model scores]
    D -->|Ensemble| G[Weighted model voting]
    E --> H[CNN-BiLSTM / RCNN]
    F --> H
    G --> H
    H --> I[Recognition / Intensity / Area]

Tech stack
#

  • Language: Python
  • Deep learning: TensorFlow 2.8 and Keras, TensorFlow Addons
  • Feature extraction: PoseNet (body keypoints), OpenFace (facial action units), OpenCV (video preprocessing)
  • Ensembles and augmentation: DeepStack (weighted ensembles), time-series data augmentation
  • Data and evaluation: pandas, NumPy, scikit-learn, k-fold cross-validation
  • Workflow: Jupyter notebooks, Cookiecutter Data Science project structure, Pipenv

Results
#

TaskBest strategyAUC
Pain recognitionBody + face ensemble0.71
Intensity estimationBody only (CNN-BiLSTM)0.75
Area classificationLate fusion (RCNN)0.75

Body movement alone was the strongest signal for estimating pain intensity, and combining it with facial features helped with recognising and localising pain. This shows that skeleton data is a useful modality for automated pain assessment, both on its own and alongside facial expressions.

The dataset is private because of ethical and privacy agreements with the participants, so the repository is provided as a reference for the models and methods.

Ideas for Continuation
#

  • Graph neural networks: a skeleton is naturally a graph, and spatio-temporal graph neural networks capture the dependencies between joints better than CNN-BiLSTMs.
  • Transformers for fusion: cross-attention would let the model learn, frame by frame, whether the face or the body is the more reliable signal.
  • Edge deployment: pose estimation now runs in real time on ordinary devices, which would make live clinical feedback possible.