Skip to content

feat: do softmax for lmhead - #256

Open
HarmonyHu wants to merge 938 commits into
sophgo:masterfrom
HarmonyHu:master
Open

feat: do softmax for lmhead#256
HarmonyHu wants to merge 938 commits into
sophgo:masterfrom
HarmonyHu:master

Conversation

@HarmonyHu

Copy link
Copy Markdown
Collaborator
  • softmax for special use, no merge

Change-Id: I7c533edfe287f12c26fb0b31f0b10722a2138933

liang.wang01 and others added 30 commits August 22, 2025 18:52
- And, Xor, Or, Not

Change-Id: Ieae37cfac4f5f0957e46dfd4b4b8e430737b1bc1
-FxGraphConverter support Indexput by scatternd
-scatternd: support inplace_add fix buffer_size

Change-Id: I6aa02a8913e8bcdffe71b1f5b9e52c94afb31e7f
- multi-core mm backend not support for bm1688

Change-Id: Idf72b00dfdbc0497dce2e406cfd1652ad092a1fb
-Reshape when the C_dims is too large

Change-Id: I6dd4ccf395da6ea42ede01dfa894feb7ea2b0b08
- The TPUC_OPT_TIMEOUT environment variable can be used to set the timeout (unit:h/m/s).
- Currently only applies to model_zoo regression.
- Does not support _os_system_log.

Change-Id: I62a52277cfaa09b56b517598e4a011582524aa14
- use options_.debugger to judge whether use manual group cost

Change-Id: I56c41c70beded59ce05b6e9a872ff483665fc3b7
- share a long prompt for different questions

Change-Id: I3a458970875b7d59ae34ea7902947f91e9b38545
-suitable for bmcv uint8 postprocess

Change-Id: I8a7f83725d416eb3da315864794ee2bca487bb6e
--search_qtable support w4a8 mode

Change-Id: I8aeacd6aaa050fb9acb662138e6e5148103745d7
the scale of logic op output should be 1.0 for correct cal in next op

Change-Id: I68e249e4b0a2fe33f96dbe9fc5798f0faa93b33c
Improve EVA02 bmodel performance

Change-Id: Ib53226b54fdfc1d98a7aba534906137b096f2d67
Match eva02 block and generate the qtable

Change-Id: I2549afd122345f195d483d9ebd835337ba89e8fe
-correct stride_h, stride_w when in_h equals stride_h and in_w equals stride_w

Change-Id: I80f0402a912792411e7a94983dc3c7981dab6214
- A16MatMul inference is not correct when group_size = 0
- A16MatMul should alway have zp, and weight should be unsign

Change-Id: Iaef75a81b4f9c8c1df12ae5524c6fd56f80485b7
- this case will switch to group_size 128

Change-Id: I2f7612412c45e5b6c03778158c9c308ee21ca34c
- as tile said
Change-Id: Icaa658071b0ba9f314d46e9afbb150eb79f45667
-reset f16 or bf16 dtype in qtable to f32

Change-Id: I9345450f3aa6bcbd8688e3307319c330d08539e2
-reset non-int8 dtype in qtable to f32

Change-Id: I607a20335fc1604252cc4777ad728bcd44ec5375
- keep sign for A16MatMul

Change-Id: Ib6d4ca0ed105bb0a3aa4a77a7111660829df9a99
-multi-core op's results cant be displayed

-fix check-diff file name bug

Change-Id: I14dcad8bf3ab6f48cadc67f4e5335a70173e48a7
- support optimized cycle modeling for multi-core case in
CycleCalculator
- reset BW in codegen process when multi-core is used for bm1688

Change-Id: Iefa913ce029c797b6069e7c27c457eb1c1cd03f2
- fix wrong dtype_len in groupnorm rstd compute

Change-Id: Iad65b379498bc9f9489180bc1305c852215e264d
- test by: llm_convert.py -m /workspace/Qwen2-7B-Instruct -s 384 -q w8bf16 -g 0  -c bm1688 --out_dir qwen2

Change-Id: Ia28624603f2a2aac4407d621019d163b66ba8a71
- there is inf in calitable causing cali info lost for this tensor, and the calitable in model zoo not match transformed mlir now.
- inf is caused in former op where, it outputs inf by selection

Change-Id: I3a7fca46cd56ec070f4802c05b778b478b796abb
- accuracy should be checked and improved in the future

Change-Id: I8ed65fa1b0e9aeceb41bfb3727605bd48baaf149
BinaryConstShiftOp quit can_be_group_small_c
Change-Id: I44bb20b9faaada6fd4e2c5a4a675360076fa842d
- if w4a16 has no group, will switch to group 128

Change-Id: I4e4357fdf3221489739cae39bc3606e808bd50ad
-correct stride_h, stride_w when in_h equals stride_h and in_w equals stride_w

Change-Id: I93c758f2b5ef04f7590d013421a87a178d941860
-add w4int8 support for Conv2d and MatMul
-add test_onnx case

Change-Id: I7a2c986da46a6a6600f563d074b92837ae6a410e
- v1.22 version info

Change-Id: Ic1b1d5e45da10c5ae3c296a07d214d0d32203382
xiangzhou.ye and others added 29 commits November 26, 2025 18:31
-add pattern for lightstereo

Change-Id: I5f71c2b29f32b40c7bf7e5c403486e3604aff58c
- llm_convert.py -m /workspace/Qwen3-VL-2B-Instruct-W4A16  -s 2048 --max_input_length 1024  --quantize w4bf16  -c bm1688 --out_dir qwen3vl_2b --max_pixels 768,768+768,384+384,384

Change-Id: I784c214cbceddd4fe6fe5d694514b281a5054e78
- assign io to first, and check io address scope

Change-Id: I632916f4f77c18f81b9708bc3514d3e6b21e5f94
- test by llm_convert.py -m /workspace/Qwen3-VL-2B-Instruct-W4A16  -s 2048 --input_length_list 1536+1024+512+256  --quantize w4bf16  -c bm1688 --out_dir qwen3vl_2b_debug --max_pixels 768,768

Change-Id: If81062c8df834f1f1a7bc0b4d1ca99c5f3a13026
- remove externs,  and uniform compiles

Change-Id: Ie9e4d0eef62cfcb4eaee0f20005cfebf4557206c
- pad flatbuffers to make weights to be 4K aligned

Change-Id: Ie3f8746564cc87247038beefa0c30b3522eda151
- set head=1 if permute_dims=6 & order[2]=3 && order[3]=2
- current implemet not support these cases

Change-Id: Ibe29c894de10305da8a14addcb86a733483ca2c5
- fix when multi-subnets, use the lino which has operands

Change-Id: I657dabaaa1968a483ae2fb21d0fb52b0ec7c91a4
- set to use f32

Change-Id: I350ddb528c5d9da9356186839c0ea62fa55399bd
- only resnet50 pass

Change-Id: I3dbe4ef6bb7733e601a313acd178e6322cf498d6
- bm1690e mm2 f8 support dorq

Change-Id: I2a616dfa259bfc46e39c34b1a56313782da592ae
- usage: bmodel_checker.py context_dir tpu_fixed.npz --dump_mode out_fixed

Change-Id: I9ea2e80d30c2fb67f2572a1036fb5665612340d9
- if not 4K aligned, will pad zeros

Change-Id: Idee90670575b0df1bc50e5a86f50715ea471b607
- model_tool --info xxx.bmodel, can check whether bmodel is 4k aligned
- model_tool --refresh xxx.bmodel, can update bmodel to 4k aligned

Change-Id: Iafaeec76deeb5619b26ec3b7ae73d7286684fe0b
-fix LSTM cv184x int8 bug

Change-Id: Ia6edd1b20f076af8c2c123bf824966353c5fb6f5
- tpulang support bm1688 num_core=2
Change-Id: Ia993e6158a91dc5ad2c073122c43e3695d85690f
-support time_fixed_subnet use new json
-add test in test_mlir_cut.sh

Change-Id: I21e39ed8a30366bad0c14e84e8d9b87b1df7dff6
-

Change-Id: I66e358a136f43acf9a6b8f5f29ba3821109f3f3f
- yolov8 needs rank = 3 opd bug got 4

Change-Id: Icf92266d335abbc5318f94418d939f1e760192b8
- disable the pass

Change-Id: Ib8ae433ab8e7c8b215548d938dbe98647b25df90
- change id of runtime: 154687

Change-Id: I204ce5e586fac23b303ac720a0fc38182df5de74
- add TPU Ar Op to GROUP_SMALL_C

Change-Id: Ic20f1dc9a0dbaef52180f45a8977034e28f95bfb
- add out channel split to W4A16MatMulPrepare in w8a16 mode
- fix split axis error when weight transposed in W4A16MatMulPrepare
- return failure if K % group != 0 in W4A16MatMulPrepare
- fix errors when lower matmul with weight transposed to a16matmul
Change-Id: I18ca93b446ef658338e9c2651951d766d44407e2
- support lora
- lm_head use [M, K]x[N,K] = [M,N], remove [M, K]x[K,N] branch

Change-Id: Ic1dc8c1b7ec30238e0b1bb7c081d7065a4a73402
- add dynamic_programming_with_structure_detect to accelerate the
layergroup searching speed for Transformer like models

Change-Id: I60dab6c0ff64d0cad69c5bcc33a15edc2afddc3b
-supported rope mode: interleaved_pairs, contiguous_halves

Change-Id: I043faf6d6dbdb627ae95fecfdb3098cb6e9c7787
- linear use lora_rank as dim 0
- lora path use the same path as lora file
- vit rope use f32

Change-Id: Ib5424ec0d2d078ce763be22594f6e3f2bcb1b64b
- softmax for special use, no merge

Change-Id: I7c533edfe287f12c26fb0b31f0b10722a2138933
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.