Qwen
/

QwQ-32B

Text Generation

text-generation-inference

Model card Files Files and versions

hzhwcmhf commited on Mar 5

Commit

c096839

·

verified ·

1 Parent(s): d76cadc

Update README.md

Files changed (1) hide show

README.md +2 -6

README.md CHANGED Viewed

@@ -68,10 +68,6 @@ text = tokenizer.apply_chat_template(
     add_generation_prompt=True
 )
-# avoid empty thought content by forcing the model to start with "<think>\n"
-response_prefix = "<think>\n"
-text += response_prefix
 model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
 generated_ids = model.generate(
@@ -83,14 +79,14 @@ generated_ids = [
 ]
 response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
-print(response_prefix + response)
 ```
 ### Usage Guidelines
 To achieve optimal performance, we recommend the following settings:
-1. **Enforce Thoughtful Output**: Ensure the model starts with "\<think\>\n" to prevent generating empty thinking content, which can degrade output quality.
 2. **Sampling Parameters**:
    - Use Temperature=0.6 and TopP=0.95 instead of Greedy decoding to avoid endless repetitions and enhance diversity.

     add_generation_prompt=True
 )
 model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
 generated_ids = model.generate(
 ]
 response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
+print(response)
 ```
 ### Usage Guidelines
 To achieve optimal performance, we recommend the following settings:
+1. **Enforce Thoughtful Output**: Ensure the model starts with "\<think\>\n" to prevent generating empty thinking content, which can degrade output quality. If you use `apply_chat_template` and set `add_generation_prompt=True`, this is already automatically implemented, but it may cause the response to lack the \<think\> tag at the beginning. This is normal behavior.
 2. **Sampling Parameters**:
    - Use Temperature=0.6 and TopP=0.95 instead of Greedy decoding to avoid endless repetitions and enhance diversity.