Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Docs
  • Enterprise
  • Pricing

  • Log In
  • Sign Up
Zhang's picture
4 2

Zhang

Chengruidong
yangwang92's profile picture OldKingMeister's profile picture
·
  • Starmys

AI & ML interests

None yet

Organizations

None yet

authored a paper 2 months ago

Chain-of-Model Learning for Language Model

Paper • 2505.11820 • Published May 17 • 121
authored a paper 3 months ago

MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention

Paper • 2504.16083 • Published Apr 22 • 9
authored a paper 8 months ago

SCBench: A KV Cache-Centric Analysis of Long-Context Methods

Paper • 2412.10319 • Published Dec 13, 2024 • 11
authored a paper 11 months ago

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

Paper • 2409.10516 • Published Sep 16, 2024 • 44
authored 2 papers about 1 year ago

MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

Paper • 2407.02490 • Published Jul 2, 2024 • 28

Parrot: Efficient Serving of LLM-based Applications with Semantic Variable

Paper • 2405.19888 • Published May 30, 2024 • 7
authored a paper over 1 year ago

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Paper • 2404.14219 • Published Apr 22, 2024 • 257
Company
TOS Privacy About Jobs
Website
Models Datasets Spaces Pricing Docs