I discovered that the compaction feature does not work correctly. Eventhough the compaction gets the model to return something the next message continues right where the last one left of. My assumption is that the compaction correctly gets the model to summaries but subsequent request dont cut off the chat history before that.
0177.20.719.593 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 29.54.311.332 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.996 (> 0.100 thold), f_keep = 1.000
[55467] 29.54.311.928 I slot launch_slot_: id 0 | task 387 | processing task, is_child = 0
[55467] 30.02.697.707 I slot print_timing: id 0 | task 387 | n_gen = 101, tg = 23.37 t/s, tg_3s = 23.60 t/s
[55467] 30.05.736.972 I slot print_timing: id 0 | task 387 | n_gen = 194, tg = 26.37 t/s, tg_3s = 30.60 t/s
[55467] 30.08.774.428 I slot print_timing: id 0 | task 387 | n_gen = 265, tg = 25.49 t/s, tg_3s = 23.37 t/s
[55467] 30.11.814.724 I slot print_timing: id 0 | task 387 | n_gen = 316, tg = 23.51 t/s, tg_3s = 16.77 t/s
[55467] 30.12.301.236 I slot print_timing: id 0 | task 387 | prompt eval time = 4106.47 ms / 640 tokens ( 6.42 ms per token, 155.85 tokens per second)
[55467] 30.12.301.238 I slot print_timing: id 0 | task 387 | eval time = 13882.63 ms / 324 tokens ( 42.98 ms per token, 23.27 tokens per second)
[55467] 30.12.301.239 I slot print_timing: id 0 | task 387 | total time = 17989.10 ms / 964 tokens
[55467] 30.12.301.239 I slot print_timing: id 0 | task 387 | graphs reused = 383
[55467] 30.12.301.243 I slot print_timing: id 0 | task 387 | draft acceptance = 0.61696 ( 211 accepted / 342 generated), mean len = 2.85
// long context i want to compact
[55467] 30.12.304.163 I slot release: id 0 | task 387 | stop processing: n_tokens = 176122, truncated = 0
10178.12.623.354 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 30.45.948.746 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = 611151107915
[55467] 30.49.168.967 I slot launch_slot_: id 0 | task 504 | processing task, is_child = 0
[55467] 30.53.320.942 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 4171, progress = 0.56, t = 3.68 s / 1133.63 tokens per second
[55467] 30.55.275.116 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 6219, progress = 0.84, t = 5.61 s / 1107.64 tokens per second
[55467] 30.55.548.147 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 6391, progress = 0.86, t = 6.28 s / 1018.43 tokens per second
[55467] 30.56.840.228 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 7415, progress = 1.00, t = 7.19 s / 1031.60 tokens per second
[55467] 30.59.959.252 I slot print_timing: id 0 | task 504 | n_gen = 162, tg = 53.53 t/s, tg_3s = 53.86 t/s
[55467] 31.02.962.961 I slot print_timing: id 0 | task 504 | n_gen = 292, tg = 48.41 t/s, tg_3s = 43.28 t/s
[55467] 31.05.972.238 I slot print_timing: id 0 | task 504 | n_gen = 447, tg = 49.44 t/s, tg_3s = 51.51 t/s
[55467] 31.09.000.928 I slot print_timing: id 0 | task 504 | n_gen = 598, tg = 49.55 t/s, tg_3s = 49.86 t/s
[55467] 31.12.049.402 I slot print_timing: id 0 | task 504 | n_gen = 750, tg = 49.61 t/s, tg_3s = 49.86 t/s
[55467] 31.13.253.884 I slot print_timing: id 0 | task 504 | prompt eval time = 7782.71 ms / 7419 tokens ( 1.05 ms per token, 953.27 tokens per second)
[55467] 31.13.253.886 I slot print_timing: id 0 | task 504 | eval time = 16301.98 ms / 794 tokens ( 20.56 ms per token, 48.64 tokens per second)
[55467] 31.13.253.886 I slot print_timing: id 0 | task 504 | total time = 24084.69 ms / 8213 tokens
[55467] 31.13.253.887 I slot print_timing: id 0 | task 504 | graphs reused = 649
[55467] 31.13.253.891 I slot print_timing: id 0 | task 504 | draft acceptance = 0.64453 ( 524 accepted / 813 generated), mean len = 2.93
[55467] 31.13.254.085 I slot release: id 0 | task 504 | stop processing: n_tokens = 8214, truncated = 0
// returned presumably the summary
10178.39.960.509 E srv operator(): http client error: Connection handling canceled
10180.54.269.523 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 33.27.873.571 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = 611212057838
[55467] 33.28.965.816 I slot launch_slot_: id 0 | task 783 | processing task, is_child = 0
[55467] 33.33.605.682 I slot print_timing: id 0 | task 783 | n_gen = 103, tg = 26.94 t/s, tg_3s = 27.20 t/s
[55467] 33.36.643.610 I slot print_timing: id 0 | task 783 | n_gen = 197, tg = 28.72 t/s, tg_3s = 30.94 t/s
[55467] 33.39.698.431 I slot print_timing: id 0 | task 783 | n_gen = 285, tg = 28.75 t/s, tg_3s = 28.81 t/s
[55467] 33.40.674.123 I slot print_timing: id 0 | task 783 | prompt eval time = 852.88 ms / 25 tokens ( 34.12 ms per token, 29.31 tokens per second)
[55467] 33.40.674.124 I slot print_timing: id 0 | task 783 | eval time = 10855.21 ms / 316 tokens ( 34.46 ms per token, 29.02 tokens per second)
[55467] 33.40.674.125 I slot print_timing: id 0 | task 783 | total time = 11708.09 ms / 341 tokens
[55467] 33.40.674.125 I slot print_timing: id 0 | task 783 | graphs reused = 736
[55467] 33.40.674.127 I slot print_timing: id 0 | task 783 | draft acceptance = 0.85019 ( 227 accepted / 267 generated), mean len = 3.55
// continues before the compaction?
[55467] 33.40.677.165 I slot release: id 0 | task 783 | stop processing: n_tokens = 176463, truncated = 0
10181.08.589.202 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 33.42.179.481 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.998 (> 0.100 thold), f_keep = 1.000
[55467] 33.42.180.025 I slot launch_slot_: id 0 | task 875 | processing task, is_child = 0
[55467] 33.48.601.836 I slot print_timing: id 0 | task 875 | n_gen = 100, tg = 25.31 t/s, tg_3s = 25.56 t/s
[55467] 33.51.663.108 I slot print_timing: id 0 | task 875 | n_gen = 192, tg = 27.39 t/s, tg_3s = 30.05 t/s
[55467] 33.54.717.522 I slot print_timing: id 0 | task 875 | n_gen = 291, tg = 28.92 t/s, tg_3s = 32.41 t/s
Thank you for your time and consideration.
Hello,
I discovered that the compaction feature does not work correctly. Eventhough the compaction gets the model to return something the next message continues right where the last one left of. My assumption is that the compaction correctly gets the model to summaries but subsequent request dont cut off the chat history before that.
Thank you for your time and consideration.
Best regards,
Joshua Krimmer