1st Place Winner - SemEval 2026 Subtask C

YoungDSMLKZ at SemEval-2026 Task 13: MIL-UniXcoder with Meta-Stacking and Handcrafted Features for AI-Generated Code Detection

Yeraly Y. Gainulla1 Aqbobek Lyceum, Kazakhstan Agzam R. Shamsadinov2 National School of Physics & Math, Kazakhstan

Abstract

We propose and validate a multi-view ensemble framework for 4-class AI-generated code detection (Human, AI, Hybrid, Adversarial) in realistic long-form repositories. Our system, Team YoungDSMLKZ, ranked 1st out of 50+ teams in SemEval-2026 Task 13 Subtask C with a macro F1 of 0.7855 (+5.2 over runner-up).

The framework combines: (i) a Dynamic Multiple Instance Learning (MIL) pipeline over UniXcoder chunks for $O(N)$-scalable long-context detection, (ii) transformer-based meta-stacking (UniXcoder and ModernBERT), and (iii) an XGBoost classifier on 200+ handcrafted stylometric features. Evidence localization analysis shows that 62.4% of decisive AI-detection signals reside beyond the standard 512-token window, validating the MIL design.

Multi-View Architecture

Our system splits diagnostic analysis into three specialized streams fused via an asymmetric class-routing rule.

Dynamic MIL Stream

Breaks long files into overlapping 512-token chunks, processes each with UniXcoder, and pools embeddings into a CatBoost classifier. Effectively defeats the "truncation trap".

O(N) Scaling Max-Pooling CatBoost Head

Meta-Stacking Stream

Uses ModernBERT (1024 tokens) & UniXcoder representations with Level-1 classifiers, fused via Level-2 Meta-CatBoost with stratified OOF training to avoid leakage.

ModernBERT 1024 Level-2 Stacking Leakage-Proof

Classical ML Stream

Extracts 200+ statistical and formatting markers (AST depth, whitespace uniformity, entropy, comments ratio) fed directly into an XGBoost classifier.

Stylometrics AST Metrics XGBoost

Interactive Stylometric Analyzer

Select an example below or paste your own code to see real-time stylometric extraction and multi-stream prediction in action.

Source Code Inspector

*Analysis computes client-side metrics similar to our XGBoost stream.

Diagnostic Metrics
Char Entropy
0.000
Comment %
0.0%
Blank Line %
0.0%
Obfuscation Index
0.00
Ensemble Prediction
Human 0%
AI Generated 0%
Hybrid Code 0%
Adversarial / Obfuscated 0%
Routing Decision Analysis

Select a code pattern and click "Inspect & Analyze" to trace how our class-routing decisions are executed.

Experimental Evaluation

Deep dive into validation scores, ablation trajectories, and long-context performance.

Ablation Waterfall

Incremental Ablation on Official Test set. Starting from a pure XGBoost baseline (0.5243), incorporating Multiple Instance Learning (+0.1088) and Meta-Stacking (+0.0664) culminates in the winning score of 0.7855.

Language Feature Importance

Feature Importance across Programming Languages. Universal traits like comment_ratio and blank_ratio provide persistent signals, while structural metrics peak for modern languages like PHP and C#.

Long Context Analysis

Long-Context Evaluation: (Left) Per-class accuracy improvement of 1024-token over 512-token transformers grows with file length. (Right) 62.4% of AI detection signals reside deep beyond the standard 512-token boundary.

Confusion Matrix

Normalized Validation Confusion Matrix. Highly consistent Human code achieves 0.98 accuracy. The hardest class to distinguish is Adversarial, which often mimics human distributions.

BibTeX & Source Citation

Cite our paper if you use our code, models, or find our architecture helpful.

@inproceedings{YoungDSMLKZsemeval2026task13,
    title={{YoungDSMLKZ} at SemEval-2026 Task 13: MIL-UniXcoder with Meta-Stacking and Handcrafted Features for AI-Generated Code Detection},
    author={Gainulla, Yeraly and Shamsadinov, Agzam},
    booktitle = {Proceedings of the 20th International Workshop on Semantic Evaluation (SemEval-2026)},
    month = jun,
    year = "2026",
    address = "San Diego, USA",
    publisher = {Association for Computational Linguistics}
}