Public work
Talks, open-source software, datasets, and demonstrations.
Talks & invited lectures
- 2025
Open source & datasets
- claude-academic-prose — A Claude Code skill that enforces clear, evaluable scientific writing: a plain-empirical voice profile, rules that strip hype and rhetorical flourish, and structural rules that make every claim traceable to its evidence.
- BitBudget Benchmark — An open embedding-compression benchmark extending my PhD learning-to-hash line; single-bit codes with re-ranking stay lossless at 32× compression.
- The Corpus is the Model — Interactive image annotation demo based on my PhD research.HF Spaces
- arxiv-acronym-gen — 10K+ acronym/expansion pairs for training LLMs.HF Datasets
- arxiv-tagged-papers — Semantic keyword tags for 3M+ arXiv papers.HF Datasets
- SpamT5 — LLM-powered email spam detection (FinLLM @ IJCAI 2023).
- CodeQUEST — LLM-powered code-quality evaluation/improvement (ISSREW 2025).
- ai-prepline — Text preprocessing pipelines for downstream NLP.
- DeepLPF — CVPR 2020 image-enhancement codebase.
- CURL — ICPR 2020 neural curve layers.
- SIDGAN — ECCV 2020 low-light video enhancement.
- SKLCRM — Sparse kernel relevance model (Best Paper, ICMR 2014).
- CryptoPanda — Cryptocurrency scanner backtest.
- Satire Classifier — Naive Bayes NLP demo.