Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

we really need LLM trained on AST, instead of token, is there any research on this?


ASTrust: Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations

https://arxiv.org/abs/2407.08983

AST-T5: Structure-Aware Pretraining for Code Generation and Understanding

https://arxiv.org/abs/2401.03003

CodeGRAG: Bridging the Gap between Natural Language and Programming Language via Graphical Retrieval Augmented Generation

https://arxiv.org/abs/2405.02355


The downside is that you need to properly preprocess code, have less non-code Training Data, and can not adapt easily to new programming languages




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: