[{"data":1,"prerenderedAt":223},["ShallowReactive",2],{"blog:\u002Fblog\u002Fgemini-3-7-multiple-witnesses":3,"site":205},{"id":4,"title":5,"author":6,"body":7,"ctaLine":186,"ctaTitle":186,"date":187,"description":188,"extension":189,"heroAlt":190,"heroImage":191,"meta":192,"navigation":193,"ogImage":194,"path":195,"seo":196,"stem":197,"tags":198,"updated":203,"__hash__":204},"blog\u002Fblog\u002Fgemini-3-7-multiple-witnesses.md","Visual analysis in BitterClip just got a lot faster","Michael Ruescher, Founder",{"type":8,"value":9,"toc":179},"minimark",[10,14,17,20,23,26,29,32,37,40,51,54,61,64,67,73,76,82,86,89,92,154,157,160,164,167,170,173,176],[11,12,13],"p",{},"Visual analysis in BitterClip just got a lot faster.",[11,15,16],{},"We upgraded our main visual model from Gemini 3.6 to Gemini 3.7. In matched\ntesting, median model response time fell from 27.8 seconds to 10.8 seconds. That\nis a 61.1% lower response time. Estimated cost was 24.6% lower too.",[11,18,19],{},"We did not find a strong semantic difference between the two models in the\nfootage we reviewed. Both generally understood the scene. So we put the faster\none in front and kept Gemini 3.6 as the fallback.",[11,21,22],{},"The route there was less direct. Our first Gemini 3.7 production test went\nbackward in time.",[11,24,25],{},"The model said a visual event began at 146 seconds and ended at 20.8 seconds. If\nyou were looking for a reason not to upgrade, there it was.",[11,27,28],{},"Then Gemini 3.6 did the same kind of thing.",[11,30,31],{},"That was the useful moment. We had blamed the new model before looking closely\nat the question we were asking it.",[33,34,36],"h2",{"id":35},"we-were-asking-for-too-much","We were asking for too much",[11,38,39],{},"BitterClip uses Gemini to understand what is happening in a video and when it\nhappens. Our old response format asked for a precise start and end for every\nvisual observation:",[41,42,48],"pre",{"className":43,"code":45,"language":46,"meta":47},[44],"language-text","start_seconds: 146.0\nend_seconds:    20.8\n","text","",[49,50,45],"code",{"__ignoreMap":47},[11,52,53],{},"The numbers were individually valid. Together, they were impossible.",[11,55,56],{},[57,58],"img",{"alt":59,"src":60},"A sanitized model response shows an impossible interval running from 146 seconds back to 20.8 seconds.","\u002Fimages\u002Fblog\u002Fgemini-3-7-benchmark\u002Ftimestamp-reversal.svg",[11,62,63],{},"The deeper problem was that video often does not contain a clean ending. A\ncamera can show someone beginning a movement, then turn away while the movement\ncontinues. Asking the model for an exact duration invites it to guess.",[11,65,66],{},"So we tried a smaller question: where can you actually see the thing?",[41,68,71],{"className":69,"code":70,"language":46,"meta":47},[44],"at_seconds: 146.0\n",[49,72,70],{"__ignoreMap":47},[11,74,75],{},"That point is still a model claim, and BitterClip still has to check it against\nthe source. But it is a much cleaner unit of evidence. BitterClip can build a\nshort playback range around it for navigation without pretending the action\nlasted for that whole range.",[11,77,78],{},[57,79],{"alt":80,"src":81},"A model-guessed duration is replaced by one model-authored evidence point and a separate playback range.","\u002Fimages\u002Fblog\u002Fgemini-3-7-benchmark\u002Fevidence-point-contract.svg",[33,83,85],{"id":84},"what-the-testing-showed","What the testing showed",[11,87,88],{},"We ran 192 matched calls across the same two video chunks and four response\nformats. We did not find a strong semantic difference between Gemini 3.6 and\n3.7 in the material we reviewed. Both could understand the scene, and both\ncould produce bad timing under the old interval format.",[11,90,91],{},"The operational difference was clear:",[93,94,95,112],"table",{},[96,97,98],"thead",{},[99,100,101,105,109],"tr",{},[102,103,104],"th",{},"Fixed test workload",[102,106,108],{"align":107},"right","Gemini 3.6",[102,110,111],{"align":107},"Gemini 3.7",[113,114,115,127,141],"tbody",{},[99,116,117,121,124],{},[118,119,120],"td",{},"Raw temporal checks passed",[118,122,123],{"align":107},"94 \u002F 96",[118,125,126],{"align":107},"92 \u002F 96",[99,128,129,132,135],{},[118,130,131],{},"Median provider latency",[118,133,134],{"align":107},"27.8 s",[118,136,137],{"align":107},[138,139,140],"strong",{},"10.8 s",[99,142,143,146,149],{},[118,144,145],{},"Estimated provider cost",[118,147,148],{"align":107},"$2.75",[118,150,151],{"align":107},[138,152,153],{},"$2.08",[11,155,156],{},"Google charged both models at the same rate; 3.7 simply used fewer tokens in\nthis test. The point format also passed all 48 of its raw timing checks.",[11,158,159],{},"A larger follow-up stopped early because of a bug in our audit code, so we did\nnot use it to declare a benchmark winner. It did not change the practical read:\n3.7 was much faster, a little cheaper, and not meaningfully different in the\nfootage we reviewed.",[33,161,163],{"id":162},"we-upgraded","We upgraded",[11,165,166],{},"We ran a strict two-chunk production canary, changed one model pin, and ran one\nmore live visual-path check. All three calls used Gemini 3.7 without falling\nback or retrying.",[11,168,169],{},"Gemini 3.7 is now the primary model for BitterClip’s visual analysis. Gemini 3.6\nremains the fallback. If wider use exposes a problem, we can put 3.6 back in\nfront with the same small pin change.",[11,171,172],{},"The point-format work remains separate from the model upgrade. The useful result\nfor people using BitterClip today is simpler: visual analysis now spends much\nless time waiting on the model. We upgraded, kept a clean way back, and moved\non.",[174,175],"hr",{},[11,177,178],{},"The comparison used 96 calls per model on two fixed silent video chunks. The\ncost figures are rate-card estimates, not a reconciled provider bill. This was a\nfocused product test, not a general benchmark of video understanding.",{"title":47,"searchDepth":180,"depth":180,"links":181},3,[182,184,185],{"id":35,"depth":183,"text":36},2,{"id":84,"depth":183,"text":85},{"id":162,"depth":183,"text":163},null,"2026-08-14","Gemini 3.7 cut median model response time by 61% in our video tests, with no strong semantic difference from 3.6. So we upgraded and kept 3.6 as the fallback.","md","BitterClip visual analysis comparison showing median model response time falling from 27.8 seconds with Gemini 3.6 to 10.8 seconds with Gemini 3.7.","\u002Fimages\u002Fblog\u002Fgemini-3-7-benchmark\u002Fgemini-3-7-benchmark-chart.svg",{},true,"\u002Fimages\u002Fblog\u002Fgemini-3-7-benchmark\u002Fgemini-3-7-benchmark-og.png","\u002Fblog\u002Fgemini-3-7-multiple-witnesses",{"title":5,"description":188},"blog\u002Fgemini-3-7-multiple-witnesses",[199,200,201,202],"Gemini","visual understanding","benchmarks","engineering","2026-08-16","dwi7iKkE-aZ8oTT7lI-QDSa6osyNGEljVbTcLlSiVdQ",{"id":206,"app_brand_name":207,"app_origin":208,"brand_name":209,"company_support_url":210,"company_url":211,"extension":212,"intended_site_origin":213,"mcp_resource_url":214,"meta":215,"og_image_default":216,"pricing_url":217,"signup_url":218,"site_origin":219,"stem":220,"support_email":221,"__hash__":222},"site\u002F_data\u002Fsite.yml","BitterClip","https:\u002F\u002Fapp.bitterclip.com","Showbrew","https:\u002F\u002Fcompany.sheetgenius.com\u002Fbitterclip\u002Fsupport\u002F","https:\u002F\u002Fcompany.sheetgenius.com\u002F","yml","https:\u002F\u002Fshowbrew.com","https:\u002F\u002Fapp.bitterclip.com\u002Fmcp",{},"\u002Fimages\u002Fshowbrew-og.png","https:\u002F\u002Fbitterclip.com\u002F#pricing","https:\u002F\u002Fapp.bitterclip.com\u002Fsign_up?plan=clip","https:\u002F\u002Fbitterclip.com","_data\u002Fsite","hello@bitterclip.com","lp_DIYW8yRANwIgWUc2X8g6xSdNBj_obL52xQi7qnu0",1788747558979]