Can Huang

NA
h-index30
6papers
204citations
Novelty47%
AI Score41

6 Papers

22.8CVNov 22, 2023Code
Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer

Zhen Zhao, Jingqun Tang, Chunhui Lin et al.

Scene text recognition (STR) in the wild frequently encounters challenges when coping with domain variations, font diversity, shape deformations, etc. A straightforward solution is performing model fine-tuning tailored to a specific scenario, but it is computationally intensive and requires multiple model copies for various scenarios. Recent studies indicate that large language models (LLMs) can learn from a few demonstration examples in a training-free manner, termed "In-Context Learning" (ICL). Nevertheless, applying LLMs as a text recognizer is unacceptably resource-consuming. Moreover, our pilot experiments on LLMs show that ICL fails in STR, mainly attributed to the insufficient incorporation of contextual information from diverse samples in the training stage. To this end, we introduce E$^2$STR, a STR model trained with context-rich scene text sequences, where the sequences are generated via our proposed in-context training strategy. E$^2$STR demonstrates that a regular-sized model is sufficient to achieve effective ICL capabilities in STR. Extensive experiments show that E$^2$STR exhibits remarkable training-free adaptation in various scenarios and outperforms even the fine-tuned state-of-the-art approaches on public benchmarks. The code is released at https://github.com/bytedance/E2STR .

1.2NAAug 26, 2014
Spectral method for substantial fractional differential equations

Can Huang, Qingshuo Song, Zhimin Zhang

In this paper, a non-polynomial spectral Petrov-Galerkin method and associated collocation method for substantial fractional differential equations (FDEs) are proposed, analyzed, and tested. We extend a class of generalized Laguerre polynomials to form our basis. By a proper scaling of trial basis and test basis, our Petrov-Galerkin method results in a diagonal and thus well-conditioned linear systems for both fractional advection equation and fractional diffusion equation. In the meantime, we construct substantial fractional differential collocation matrices and provide explicit forms for both type of equations. Moreover, the proposed method allows us to adjust a parameter in basis selection according to different given data to maximize the convergence rate. This fact has been proved in our error analysis and confirmed in our numerical experiments.

1.2NAJan 24, 2018
An accurate spectral method for Maxwell equations in Cole-Cole dispersive media

Can Huang, Li-lian Wang

In this paper, we propose an accurate numerical means built upon a spectral-Galerkin method in spatial discretization and an enriched multi-step spectral-collocation approach in temporal direction, for Maxwell equations in Cole-Cole dispersive media in two-dimensional setting. Our starting point is to derive a new model involving only one unknown field from the original model with three unknown fields: electric, magnetic fields and the induced electric polarisation (described by a global temporal convolution of the electric field). This results in a second-order integral-differential equation with a weakly singular integral kernel expressed by the Mittag-Lefler (ML) function. The most interesting but challenging issue resides in how to efficiently deal with the singularity in time induced by the ML function which is an infinite series of singular power functions with different nature. With this in mind, we introduce a spectral-Galerkin method using Fourier-like basis functions for spatial discretization, leading to a sequence of decoupled temporal integral-differential equations (IDE) with the same weakly singular kernel involving the ML function as the original two-dimensional problem. With a careful study of the regularity of IDE, we incorporate several leading singular terms into the numerical scheme and approximate much regular part of the solution. Then we solve to IDE by a multi-step well-conditioned collocation scheme together with mapping technique to increase the accuracy and enhance the resolution. We show such an enriched collocation method is convergent and accurate. % analysis of the scheme is carried out.

30.3CVMay 20, 2024Code
MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Jingqun Tang, Qi Liu, Yongjie Ye et al.

Text-Centric Visual Question Answering (TEC-VQA) in its proper format not only facilitates human-machine interaction in text-centric visual environments but also serves as a de facto gold proxy to evaluate AI models in the domain of text-centric scene understanding. Nonetheless, most existing TEC-VQA benchmarks have focused on high-resource languages like English and Chinese. Despite pioneering works to expand multilingual QA pairs in non-text-centric VQA datasets through translation engines, the translation-based protocol encounters a substantial "visual-textual misalignment" problem when applied to TEC-VQA. Specifically, it prioritizes the text in question-answer pairs while disregarding the visual text present in images. Moreover, it fails to address complexities related to nuanced meaning, contextual distortion, language bias, and question-type diversity. In this work, we tackle multilingual TEC-VQA by introducing MTVQA, the first benchmark featuring high-quality human expert annotations across 9 diverse languages, consisting of 6,778 question-answer pairs across 2,116 images. Further, by comprehensively evaluating numerous state-of-the-art Multimodal Large Language Models~(MLLMs), including Qwen2-VL, GPT-4o, GPT-4V, Claude3, and Gemini, on the MTVQA benchmark, it is evident that there is still a large room for performance improvement (Qwen2-VL scoring 30.9 versus 79.7 for human performance), underscoring the value of MTVQA. Additionally, we supply multilingual training data within the MTVQA dataset, demonstrating that straightforward fine-tuning with this data can substantially enhance multilingual TEC-VQA performance. We aspire that MTVQA will offer the research community fresh insights and stimulate further exploration in multilingual visual text comprehension. The project homepage is available at https://bytedance.github.io/MTVQA/.

16.3CLSep 10, 2025
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

Haiyang Yu, Yuchuan Wu, Fan Shi et al.

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while current Vision-Language Models (VLMs) struggle with their visual and linguistic complexity. Existing document benchmarks focus on English printed texts or simplified Chinese, leaving a gap for evaluating VLMs on ancient Chinese documents. To address this, we present AncientDoc, the first benchmark for Chinese ancient documents, designed to assess VLMs from OCR to knowledge reasoning. AncientDoc includes five tasks (page-level OCR, vernacular translation, reasoning-based QA, knowledge-based QA, linguistic variant QA) and covers 14 document types, over 100 books, and about 3,000 pages. Based on AncientDoc, we evaluate mainstream VLMs using multiple metrics, supplemented by a human-aligned large language model for scoring.

1.2NAMar 27, 2015
Well-Conditioned Fractional Collocation Methods Using Fractional Birkhoff Interpolation Basis

Yujian Jiao, Li-Lian Wang, Can Huang

The purpose of this paper is twofold. Firstly, we provide explicit and compact formulas for computing both Caputo and (modified) Riemann-Liouville (RL) fractional pseudospectral differentiation matrices (F-PSDMs) of any order at general Jacobi-Gauss-Lobatto (JGL) points. We show that in the Caputo case, it suffices to compute F-PSDM of order $μ\in (0,1)$ to compute that of any order $k+μ$ with integer $k\ge 0,$ while in the modified RL case, it is only necessary to evaluate a fractional integral matrix of order $μ\in (0,1).$ Secondly, we introduce suitable fractional JGL Birkhoff interpolation problems leading to new interpolation polynomial basis functions with remarkable properties: (i) the matrix generated from the new basis yields the exact inverse of F-PSDM at "interior" JGL points; (ii) the matrix of the highest fractional derivative in a collocation scheme under the new basis is diagonal; and (iii) the resulted linear system is well-conditioned in the Caputo case, while in the modified RL case, the eigenvalues of the coefficient matrix are highly concentrated. In both cases, the linear systems of the collocation schemes using the new basis can solved by an iterative solver within a few iterations. Notably, the inverse can be computed in a very stable manner, so this offers optimal preconditioners for usual fractional collocation methods for fractional differential equations (FDEs). It is also noteworthy that the choice of certain special JGL points with parameters related to the order of the equations can ease the implementation. We highlight that the use of the Bateman's fractional integral formulas and fast transforms between Jacobi polynomials with different parameters, are essential for our algorithm development.