Summary: Automated protein function annotation remains challenging as sequence databases outpace curated labels and homology-based transfer fails for proteins lacking close relatives. We present GOlien, a scalable, composition-based method that annotates protein sequences using a Shannon-entropy k-mer model. Built from the CAFA3-based training split and evaluated on the held-out validation split, GOlien achieves micro-averaged of 0.7317 (Biological Process), 0.7610 (Cellular Component) and 0.8295 (Molecular Function). Our tool offers a practical, complementary alternative to existing pipelines and extending annotation coverage for proteins. Availability and Implementation: The command line tool to submit FASTA files and retrieve predicted GO term annotations with the corresponding documentation is available at: https://github.com/probalytiq/golien-tool
Laczko, L., Pek, D., Kovacs, B. L.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 11
- Comments 0
