Instructions to use mlx-community/gemma-3n-E2B-6bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use mlx-community/gemma-3n-E2B-6bit with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="mlx-community/gemma-3n-E2B-6bit")

# Load model directly
from transformers import AutoProcessor, AutoModelForImageTextToText

processor = AutoProcessor.from_pretrained("mlx-community/gemma-3n-E2B-6bit")
model = AutoModelForImageTextToText.from_pretrained("mlx-community/gemma-3n-E2B-6bit")

MLX

How to use mlx-community/gemma-3n-E2B-6bit with MLX:

# Make sure mlx-vlm is installed
# pip install --upgrade mlx-vlm

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

# Load the model
model, processor = load("mlx-community/gemma-3n-E2B-6bit")
config = load_config("mlx-community/gemma-3n-E2B-6bit")

# Prepare input
image = ["http://images.cocodataset.org/val2017/000000039769.jpg"]
prompt = "Describe this image."

# Apply chat template
formatted_prompt = apply_chat_template(
    processor, config, prompt, num_images=1
)

# Generate output
output = generate(model, processor, formatted_prompt, image)
print(output)

Notebooks
Google Colab
Kaggle
Local Apps
LM Studio

vLLM

How to use mlx-community/gemma-3n-E2B-6bit with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "mlx-community/gemma-3n-E2B-6bit"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "mlx-community/gemma-3n-E2B-6bit",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker

docker model run hf.co/mlx-community/gemma-3n-E2B-6bit

SGLang

How to use mlx-community/gemma-3n-E2B-6bit with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "mlx-community/gemma-3n-E2B-6bit" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "mlx-community/gemma-3n-E2B-6bit",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "mlx-community/gemma-3n-E2B-6bit" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "mlx-community/gemma-3n-E2B-6bit",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Docker Model Runner
How to use mlx-community/gemma-3n-E2B-6bit with Docker Model Runner:
```
docker model run hf.co/mlx-community/gemma-3n-E2B-6bit
```

prince-canuma commited on Jul 12, 2025

Commit

34ecf7b

verified ·

1 Parent(s): 6cdea72

Upload folder using huggingface_hub

Browse files

Files changed (6) hide show

README.md +1 -1
config.json +2 -6
generation_config.json +1 -1
model-00001-of-00002.safetensors +2 -2
model.safetensors.index.json +3 -2
preprocessor_config.json +1 -1

README.md CHANGED Viewed

@@ -17,7 +17,7 @@ tags:
 ---
 # mlx-community/gemma-3n-E2B-6bit
-This model was converted to MLX format from [`google/gemma-3n-E2B`]() using mlx-vlm version **0.3.0**.
 Refer to the [original model card](https://huggingface.co/google/gemma-3n-E2B) for more details on the model.
 ## Use with mlx

 ---
 # mlx-community/gemma-3n-E2B-6bit
+This model was converted to MLX format from [`google/gemma-3n-E2B`]() using mlx-vlm version **0.3.1**.
 Refer to the [original model card](https://huggingface.co/google/gemma-3n-E2B) for more details on the model.
 ## Use with mlx

config.json CHANGED Viewed

@@ -63,9 +63,7 @@
         "task_specific_params": null,
         "problem_type": null,
         "_name_or_path": "",
-        "conf_positional_bias_size": 256,
         "model_type": "gemma3n_audio",
-        "sscp_conv_eps": 0.001,
         "input_feat_size": 128,
         "hidden_size": 1536,
         "rms_norm_eps": 1e-06,
@@ -3744,9 +3742,7 @@
         "task_specific_params": null,
         "problem_type": null,
         "_name_or_path": "",
-        "altup_lr_multiplier": 1.0,
         "model_type": "gemma3n_text",
-        "query_pre_attn_scalar": 256,
         "vocab_size": 262400,
         "vocab_size_per_layer_input": 262144,
         "max_position_embeddings": 32768,
@@ -3878,7 +3874,7 @@
     "top_k": 50,
     "top_p": 1.0,
     "torchscript": false,
-    "transformers_version": "4.53.1",
     "typical_p": 1.0,
     "use_bfloat16": false,
     "vision_config": {
@@ -3940,7 +3936,7 @@
         "model_type": "gemma3n_vision",
         "num_classes": 2,
         "initializer_range": 0.02,
-        "do_pooling": true,
         "model_args": null,
         "architecture": "mobilenetv5_300m_enc",
         "hidden_size": 2048,

         "task_specific_params": null,
         "problem_type": null,
         "_name_or_path": "",
         "model_type": "gemma3n_audio",
         "input_feat_size": 128,
         "hidden_size": 1536,
         "rms_norm_eps": 1e-06,
         "task_specific_params": null,
         "problem_type": null,
         "_name_or_path": "",
         "model_type": "gemma3n_text",
         "vocab_size": 262400,
         "vocab_size_per_layer_input": 262144,
         "max_position_embeddings": 32768,
     "top_k": 50,
     "top_p": 1.0,
     "torchscript": false,
+    "transformers_version": "4.53.2",
     "typical_p": 1.0,
     "use_bfloat16": false,
     "vision_config": {
         "model_type": "gemma3n_vision",
         "num_classes": 2,
         "initializer_range": 0.02,
+        "do_pooling": false,
         "model_args": null,
         "architecture": "mobilenetv5_300m_enc",
         "hidden_size": 2048,

generation_config.json CHANGED Viewed

@@ -6,5 +6,5 @@
   "pad_token_id": 0,
   "top_k": 64,
   "top_p": 0.95,
-  "transformers_version": "4.53.0.dev0"
 }

   "pad_token_id": 0,
   "top_k": 64,
   "top_p": 0.95,
+  "transformers_version": "4.54.0.dev0"
 }

model-00001-of-00002.safetensors CHANGED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:46da528f3424b4c2aa31d09982b202eb5721f17dd277265fcc211623d94b14fe
-size 5368407299

 version https://git-lfs.github.com/spec/v1
+oid sha256:b25ddd00a44bf592444642f18735f7e6e7c4aa880af9d2ca71700990091c3668
+size 5368407542

model.safetensors.index.json CHANGED Viewed

@@ -1,7 +1,7 @@
 {
   "metadata": {
-    "total_parameters": 5976833408,
-    "total_size": 10878876416
   },
   "weight_map": {
     "model.audio_tower.conformer.0.attention.attn.k_proj.weight": "model-00001-of-00003.safetensors",
@@ -1553,6 +1553,7 @@
     "model.vision_tower.timm_model.blocks.3.9.layer_scale.gamma": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.blocks.3.9.norm.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.conv_stem.bn.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.conv_stem.conv.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.msfa.ffn.pw_exp.bn.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.msfa.ffn.pw_exp.conv.weight": "model-00001-of-00003.safetensors",

 {
   "metadata": {
+    "total_parameters": 5976833472,
+    "total_size": 10878876544
   },
   "weight_map": {
     "model.audio_tower.conformer.0.attention.attn.k_proj.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.blocks.3.9.layer_scale.gamma": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.blocks.3.9.norm.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.conv_stem.bn.weight": "model-00001-of-00003.safetensors",
+    "model.vision_tower.timm_model.conv_stem.conv.bias": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.conv_stem.conv.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.msfa.ffn.pw_exp.bn.weight": "model-00001-of-00003.safetensors",
     "model.vision_tower.timm_model.msfa.ffn.pw_exp.conv.weight": "model-00001-of-00003.safetensors",

preprocessor_config.json CHANGED Viewed

@@ -41,7 +41,7 @@
   "processor_class": "Gemma3nProcessor",
   "resample": 2,
   "rescale_factor": 0.00392156862745098,
-  "return_attention_mask": false,
   "return_tensors": null,
   "sampling_rate": 16000,
   "size": {

   "processor_class": "Gemma3nProcessor",
   "resample": 2,
   "rescale_factor": 0.00392156862745098,
+  "return_attention_mask": true,
   "return_tensors": null,
   "sampling_rate": 16000,
   "size": {