-
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
Paper • 2408.13233 • Published • 22 -
Heterogeneous Multi-task Learning with Expert Diversity
Paper • 2106.10595 • Published • 2 -
Residual Mixture of Experts
Paper • 2204.09636 • Published • 1 -
Language-Routing Mixture of Experts for Multilingual and Code-Switching Speech Recognition
Paper • 2307.05956 • Published • 1
Hazem Essam
hazemessam
AI & ML interests
Protein Language Modeling and Natural Language Processing
Recent Activity
updated a model 4 days ago
hazemessam/ares-softmoe-4b-l2-consecutive-225K updated a model 5 days ago
HazemLab/ares-softmoe-4b-l2-consecutive-225K updated a model 5 days ago
HazemLab/ares-softmoe-4b-l2-consecutive-150K