ChatPaper.ai
打开菜单
首页
每日论文
工作台
定价
账户
🇨🇳
中文简体
Loading...
•
•
•
•
•
•
•
•
•
•
AI研究论文每日精选
每日精选AI研究论文及翻译
November 18th, 2024
LLaVA-o1:让视觉语言模型逐步推理
LLaVA-o1: Let Vision Language Models Reason Step-by-Step
Guowei Xu, Peng Jin, Li Hao, Yibing Song, Lichao Sun, Li Yuan
•
Nov 15, 2024
•
105
7
高斯任意性:用于三维生成的交互式点云潜在扩散
GaussianAnything: Interactive Point Cloud Latent Diffusion for 3D Generation
Yushi Lan, Shangchen Zhou, Zhaoyang Lyu, Fangzhou Hong, Shuai Yang, Bo Dai, Xingang Pan, Chen Change Loy
•
Nov 12, 2024
•
21
6
通过硬绑定和软细化实现区域感知的文本到图像生成
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
Zhennan Chen, Yajie Li, Haofan Wang, Zhibo Chen, Zhengkai Jiang, Jun Li, Qian Wang, Jian Yang, Ying Tai
•
Nov 10, 2024
•
34
6
GUI代理的曙光:与克劳德3.5计算机的初步案例研究
The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use
Siyuan Hu, Mingyu Ouyang, Difei Gao, Mike Zheng Shou
•
Nov 15, 2024
•
29
2
Xmodel-1.5:一个规模为10亿的多语言LLM
Xmodel-1.5: An 1B-scale Multilingual LLM
Wang Qun, Liu Yang, Lin Qingquan, Jiang Ling
•
Nov 15, 2024
•
14
2
MARS:释放方差减少在训练大型模型中的力量
MARS: Unleashing the Power of Variance Reduction for Training Large Models
Huizhuo Yuan, Yifeng Liu, Shuang Wu, Xun Zhou, Quanquan Gu
•
Nov 15, 2024
•
13
2
对视频进行时间定位,就像翻阅漫画一样。
Number it: Temporal Grounding Videos like Flipping Manga
Yongliang Wu, Xinting Hu, Yuyang Sun, Yizhou Zhou, Wenbo Zhu, Fengyun Rao, Bernt Schiele, Xu Yang
•
Nov 15, 2024
•
12
2