We propose and validate a multi-view ensemble framework for 4-class AI-generated code detection (Human, AI, Hybrid, Adversarial) in realistic long-form repositories. Our system, Team YoungDSMLKZ, ranked 1st out of 50+ teams in SemEval-2026 Task 13 Subtask C with a macro F1 of 0.7855 (+5.2 over runner-up).
The framework combines: (i) a Dynamic Multiple Instance Learning (MIL) pipeline over UniXcoder chunks for $O(N)$-scalable long-context detection, (ii) transformer-based meta-stacking (UniXcoder and ModernBERT), and (iii) an XGBoost classifier on 200+ handcrafted stylometric features. Evidence localization analysis shows that 62.4% of decisive AI-detection signals reside beyond the standard 512-token window, validating the MIL design.
Our system splits diagnostic analysis into three specialized streams fused via an asymmetric class-routing rule.
Breaks long files into overlapping 512-token chunks, processes each with UniXcoder, and pools embeddings into a CatBoost classifier. Effectively defeats the "truncation trap".
Uses ModernBERT (1024 tokens) & UniXcoder representations with Level-1 classifiers, fused via Level-2 Meta-CatBoost with stratified OOF training to avoid leakage.
Extracts 200+ statistical and formatting markers (AST depth, whitespace uniformity, entropy, comments ratio) fed directly into an XGBoost classifier.
Select an example below or paste your own code to see real-time stylometric extraction and multi-stream prediction in action.
*Analysis computes client-side metrics similar to our XGBoost stream.
Select a code pattern and click "Inspect & Analyze" to trace how our class-routing decisions are executed.
Deep dive into validation scores, ablation trajectories, and long-context performance.
Incremental Ablation on Official Test set. Starting from a pure XGBoost baseline (0.5243), incorporating Multiple Instance Learning (+0.1088) and Meta-Stacking (+0.0664) culminates in the winning score of 0.7855.
Feature Importance across Programming Languages. Universal traits like comment_ratio and blank_ratio provide persistent signals, while structural metrics peak for modern languages like PHP and C#.
Long-Context Evaluation: (Left) Per-class accuracy improvement of 1024-token over 512-token transformers grows with file length. (Right) 62.4% of AI detection signals reside deep beyond the standard 512-token boundary.
Normalized Validation Confusion Matrix. Highly consistent Human code achieves 0.98 accuracy. The hardest class to distinguish is Adversarial, which often mimics human distributions.
Cite our paper if you use our code, models, or find our architecture helpful.
@inproceedings{YoungDSMLKZsemeval2026task13,
title={{YoungDSMLKZ} at SemEval-2026 Task 13: MIL-UniXcoder with Meta-Stacking and Handcrafted Features for AI-Generated Code Detection},
author={Gainulla, Yeraly and Shamsadinov, Agzam},
booktitle = {Proceedings of the 20th International Workshop on Semantic Evaluation (SemEval-2026)},
month = jun,
year = "2026",
address = "San Diego, USA",
publisher = {Association for Computational Linguistics}
}