Accurate and robust pose estimation is essential for automated livestock monitoring, particularly in group-housed pigs where occlusions, animal overlap, and variable orientations complicate keypoint detection. This study evaluated the effects of keypoint configuration and neural network architecture on pose estimation performance in group-housed fattening pigs. Five skeletal configurations were first compared using ResNet-50 in DeepLabCut: a standard 7-keypoint model and four extended configurations including additional ear, dorsal body, eye, or tail-related landmarks. BODY5P, which included five anatomically distributed landmarks along the dorsal body axis, achieved the best overall test performance, with the highest mAP (75.12 ± 1.42%) and the lowest RMSE (7.43 ± 0.43 px), suggesting that consistently visible dorsal landmarks improve robustness under group-housing conditions. Five neural network architectures were then compared using a larger annotated dataset and repeated train-test evaluation. HRNet-W32 achieved the highest accuracy (test mAP: 92.52%; RMSE: 7.93 ± 0.36 px), whereas lighter DLCRNet architectures trained faster and showed competitive test performance, indicating their potential for resource-constrained or frequently retrained on-farm applications. Temporal analysis of high-confidence predictions from an independent 15-min video showed stable tracking for most dorsal landmarks, while the snout had the highest dropout rate (∼46%). Exploratory spatial analysis showed that pose outputs can support occupancy mapping and tail-snout proximity analysis, although proximity events should be interpreted as candidate interaction indicators rather than confirmed tail-biting events. Overall, these results highlight the importance of balancing anatomical relevance, landmark visibility, model accuracy, and computational efficiency when designing pose-estimation pipelines for pig behaviour monitoring.